
If your child has ever spent an hour glued to a beautifully made learning app and come away no better at the actual skill, you already understand the gap at the center of this story. On September 17, Google Research announced a tool that lets a teacher describe a learning objective and have an AI build a custom, interactive STEM simulation for it in seconds, complete with hints, feedback, and worked solutions. The early results look encouraging: teachers and expert reviewers rated the simulations highly for quality and usability. Here is the part the announcement is careful about and the coverage tends to skip. A tool that experts rate as good-looking and a tool that actually teaches are two different claims, and only the first one has been tested so far.
Google’s new AI lets teachers spin up polished STEM simulations in seconds, and expert raters love them. Here is the catch the headlines skip: a high quality rating measures how good a tool looks, not whether a child learned from it. That part has not been tested yet.
Google’s new AI builds a classroom simulation in seconds, but a good rating from a teacher is not the same as proof a child learned. Here are the questions parents are asking about what that means at home.
Common questions
Does a high teacher rating mean an AI learning tool will help my child learn?
What actually makes a learning tool effective?
How do I tell if a tool is teaching my child rather than only keeping them busy?
My child is struggling and I am not sure which tools to trust. What should I do at home?
Google's new AI builds slick classroom simulations and teachers rate them highly. But a high rating measures how good it looks, not whether a child learned. Nobody has tested that yet.
What Google actually announced
In a September 17, 2026 post titled ‘The future of practice,’ Google Research scientists Gal Elidan and Yael Haramaty described a system, built on Gemini, that turns a teacher’s request into an interactive learning simulation. A teacher submits an objective, the pipeline generates several candidate simulations, and automated raters plus the teacher choose the one to use. A companion academic paper on arXiv lays out the method in detail, including an automated self-check that opens the simulation and interacts with it like a student, adversarial inputs and all, before it ever reaches a classroom.
The validation is where the story turns. In a US pilot, 12 teachers each requested three interactives, 36 in all, and the blog reports ‘an average teacher rating of 8 out of 10 on the interactives’ quality.’ The paper is more precise, reporting ‘an average rating of above 8.1, corresponding to good usability,’ with ‘all simulations received ratings of 7 or above on usability.’ A separate UK expert review of 40 simulation requests, each scored by two independent teacher raters, reported an 86 percent acceptance rate. In the blog’s words, ‘overall rating is good or excellent with physics and chemistry being the most amenable to simulation creation.’
The design itself reads like a checklist of things careful teachers ask for: an introduction to prime prior knowledge, a toolbox of relevant formulas, ‘multiple levels of hints, tailored feedback reflecting on why a specific response is working or not working, and worked solutions to strengthen comprehension after the student’s own exploration.’ The researchers are also candid about what they have not done. The blog says the team ‘will be conducting UX research and field studies to evaluate learning gains and student engagement,’ and the paper calls the work a first step, naming student learning efficacy in real schools as the most important thing still to test.
the most notable need is for evaluation in the setting of real teachers in real schools, along with student learning efficacy.
Laura Lurns · Learning Success expert
The frame the science supports
Here is the quiet substitution worth catching. A rating from a teacher, however experienced, measures how good a simulation looks and how usable it feels. It does not measure whether a child who used it learned more than a child who did not. Those are different questions, and the evidence on screen-based learning tools is blunt about it: time spent, lessons completed, and how polished a thing appears are engagement and quality signals, not learning outcomes. Conflating the two is the most common move in the whole ed-tech business, and it is the framing worth resisting here.
None of that is a knock on the design, which is genuinely promising. The features Google describes, leveled goals that ask a student to manipulate variables and reason toward an answer, layered hints, and feedback that responds to why a particular solution works, line up with the categories of educational technology that learning science has associated with real teaching: tools that engage a learner at the level of their reasoning rather than only marking a final answer, and feedback that tells a student what to do next rather than flashing a red X. That is a hopeful direction. It is also, so far, a resemblance. A described feature is not the same as a feature shown to work, and plenty of apps that advertise ‘adaptive’ hints and ‘instant feedback’ never deliver the step-by-step engagement the research actually rewards.
The honest reading is the one the researchers themselves land on. The paper’s own abstract names ‘measurably effective outcomes’ as the bar that interactive learning experiences have traditionally had to clear, and then stops short of claiming this new system has cleared it. That is not a gotcha, it is the paper being careful. An independent outlet reached the same conclusion without being prompted, noting that the ratings measure perceived quality in a small initial group and not whether students learned more. Until a study puts these simulations in front of real students and measures what they learned against a fair comparison, the tool is promising and unproven at the same time, and both halves of that sentence are true.
Key takeaways
- A rating is not a result: teachers scored the simulations highly, but nobody has yet measured whether students actually learned more.
- The design is promising: leveled goals, layered hints, and next-step feedback mirror features learning science ties to tools that work.
- The real test is still ahead: Google says field studies on learning gains are coming, so watch closely what they choose to measure.
What it means for your child
This is genuinely good news for classrooms, and it deserves to be welcomed as such. A teacher who once had to accept whatever an off-the-shelf simulation happened to offer will soon be able to shape one around the exact concept a class is stuck on. That is real progress, and skepticism about the learning claim is not skepticism about the tool. Both things hold at once, which is exactly the posture a parent wants when a new technology lands in the classroom.
So keep the question simple and hold it steady: does anyone have evidence that children learned more, or only evidence that adults liked it? When a school or a company points to high ratings, high engagement, or a sleek demo, that is information about appeal, not about learning. The question to ask is whether there is an independent study with a real comparison group and a standardized outcome, rather than a rating and rather than a measure built to match what the tool happens to drill. Google has promised those field studies. The thing to watch for is not whether the studies happen but what they choose to measure.
At home, the same discipline protects your child. Engagement is easy to see and easy to trust, and a child who is busy and happy on a screen looks like a child who is learning. The truer test is whether the skill holds up away from the tool: reach for a word, a problem, or an idea when the hints and the animation are gone. That is the signal worth trusting, and it is the one no quality rating will ever hand you.
That measures perceived quality in a small initial group, not whether students learned more from the simulations.
Laura Lurns · Learning Success expert
You do not need to be a scientist to ask the one question that cuts through the hype: not ‘does it look good?’ but ‘did my child learn?’ Children are capable of real learning, and they deserve tools measured by that, not by how engaging a demo feels or how many experts nodded along. The villain here is not Google, and it is not this genuinely clever tool. It is the old, seductive framing that treats polish and high ratings as proof of learning, a habit that has sold parents shiny apps for two decades. Learning Success was built on the opposite bet, that what matters is the skill your child carries away from the screen. Our All-Access membership opens an assessment that asks about the skills your child’s learning runs on and a roadmap that names what to build first.
See what All Access gives your childIs your child struggling in school?
Get your free personalized learning roadmap
You describe what you see at home. We turn it into a plan you start this week.
- Answer 5 short questionnaires about what you already notice, 30–45 minutes at your own kitchen table
- Your child sits no test and gets no score: nothing to schedule, nothing for them to dread
- You do the answering, the AI does the writing, and a person reviews it before it reaches you
- Access all 40+ courses instantly: reading, math, focus, processing and more, with new ones added regularly
Why we use AI, plainly: it writes from a knowledge base our team maintains and audits. We work through it line by line and pull anything the evidence stops supporting. The roadmap you get on Tuesday reflects what we corrected on Monday, and a human still reads it before you do.
Your school district must evaluate your child free of charge if you ask in writing, whatever your income and whatever the outcome (US, 34 CFR 300.111 and 300.301(b)). That route takes time and answers a different question than you do. This one starts today, from what you already know.
Your answers stay yours. We do not sell your personal information, and we do not hand identifiable assessment data to outside AI companies to train their models.
A screener is a starting point, not a diagnosis. If your child might need formal accommodations (an IEP or 504 plan), or you suspect a vision, hearing or medical cause, pursue a professional evaluation too. That is the only route to those supports.
References
- Google Research — The future of practice: Enabling teachers to create learning interactives with generative UI
- arXiv:2609.20738 — Harnessing Generative UI for Education: Tailored Learning Interactives (abstract)
- arXiv:2609.20738 — Harnessing Generative UI for Education: Tailored Learning Interactives (full text)
- Superpower Daily — Google Research opens sign-ups for teacher-requested AI STEM simulation pilot



