A small group of students in Dayton, Ohio used a social-emotional learning chatbot named Jordan on their laptops during several weeks of last school year. The Alternative School at Jackson Center, part of Dayton Public Schools, serves students who would otherwise be suspended or expelled and often struggle with behavioral and academic challenges. The school paid the ed tech company SchoolAI $72,600 last academic year for multiple AI tools, including Jordan, which administrators designed in collaboration with the company and an outside consultant.
Families and students were notified that chat logs could be viewed by school employees, an external evaluator and SchoolAI staff. Jordan asks students for advice rather than directly asking how they feel, grading responses for engagement and authenticity. K, a sophomore at the school, said Jordan was sometimes frustrating but relatable, helping him solve personal problems by putting himself in Jordan’s perspective.
The Jordan Pilot
Dayton schools are running Jordan as a pilot, not a full rollout. That distinction matters. A pilot tests whether a tool works. A rollout assumes it does. The school is treating Jordan as the former, which suggests a degree of caution. But the pilot is also being evaluated by an outside consultant hired by Dayton, which adds a layer of oversight that not all districts have.
Mike Kentz, the external consultant hired by Dayton to design and evaluate the pilot, said the experiment produced enough positive evidence from chat logs and student/staff feedback to continue testing. That is the kind of outcome a pilot hopes for. But Kentz believes Jordan isn’t ready for rollout across the district because it’s difficult to measure success concretely, saying he can only measure perception.
“Every test, every experiment that I try to measure in some sort of concrete way, I can only get as far as perception,” he says, referring to how students and staff feel about it.
What the Research Says
There is no consensus on how schools should measure what works and what doesn’t with AI. A recent Stanford University review of more than 800 academic papers found extremely limited research on AI in schools, with some studies suggesting carefully designed tools show more promise than general-purpose chatbots. That review points to a gap in the science.
Robin Lake, director of the Center on Reinventing Public Education, described the situation as feeling “very Wild West in terms of the randomness.” That captures the moment. Districts are buying tools, running pilots, and collecting feedback while researchers are still figuring out what questions to ask.
The Metrics Problem
Katelyn Schoenhofer, AI specialist for Wichita Public Schools, said her district wrestles with defining success and metrics for evaluating AI use, having not reached an answer yet. Wichita schools have slowly introduced AI in controlled ways, with early grades learning about robots as pattern recognition without access to chatbots, and middle schoolers getting limited access to tools like Canva AI for image generation.
The lack of metrics is the core problem. Without agreed standards, every district is measuring success in its own way, which means comparisons across districts are nearly impossible.
The Oversight Question
The notification that chat logs could be viewed by school employees, an external evaluator and SchoolAI staff raises a question about transparency. Who gets to see the logs? Under what circumstances? For how long? These are the kinds of questions that need answering before a tool is rolled out to more classrooms.
The evaluation arrangement with Kentz involves SchoolAI, which adds a layer of complexity. Districts should ask whether the arrangement creates a conflict of interest and whether there is a way to separate the evaluation from the sale.
The Vulnerable Students
The students at Jackson Center are not typical K-12 pupils. They are students who would otherwise be suspended or expelled, and they often struggle with behavioral and academic challenges. That makes the stakes higher. A poorly designed tool could do harm. A well-designed tool could help. Either way, the decision affects young people whose records are already fragile.
Jordan’s approach — asking students for advice rather than directly asking how they feel — is a design choice. It avoids the trap of a machine pretending to understand emotions. But it also limits what the tool can measure. Engagement and authenticity are real things, but they are not the same as academic progress or social-emotional growth.
The Future of AI in Schools
The Wild West metaphor is apt. Districts are experimenting, evaluating, and adjusting as they go. Some will find tools that work. Some will find tools that don’t. The field lacks a common language for describing outcomes, and it lacks a common method for measuring them. Both need to arrive sooner rather than later.
The Stanford review’s finding that carefully designed tools show more promise than general-purpose chatbots points to a path forward. Districts should look for tools built with specific educational goals in mind, rather than general-purpose tools that are marketed as silver bullets. And they should demand research-backed evidence before committing to a purchase.
What Dayton Should Do Next
Dayton has a working pilot and a consultant who thinks it should continue. That is a reasonable position. But Kentz’s view that Jordan isn’t ready for rollout across the district is a concern worth noting. Scaling up before the measurement problem is solved would be premature.
The district should treat Kentz’s positive evidence as a green light for more testing, not a mandate for deployment. It should also work with researchers to develop clearer metrics for what success looks like.
The Bottom Line
AI is coming to schools. The question is not whether it will arrive, but how it will arrive. Dayton is doing the hard work of testing before deploying. That is better than rushing ahead. But the broader picture is worrying. The Stanford review’s finding of extremely limited research on AI in schools is a wake-up call for the entire sector.
Schools should proceed with caution. They should demand transparency from vendors. They should measure what they are doing. The Wild West era will not last forever. But it will last as long as schools keep buying tools without knowing what they are buying.
The opinion follows: SchoolAI should be held to the highest standard of transparency, and districts should measure what they are doing before they decide whether to buy.
Where the paper stands
The paper backs narrow rules against direct harm, such as forcing companies to disclose safety failures they hid, and opposes broad rules that hand the market to the incumbents. The paper also supports a light touch on AI and technology so startups are not frozen out. In Dayton’s case, the pilot design — an external consultant evaluating a pilot that administrators designed in collaboration with SchoolAI — shows the kind of caution the paper favors over a rushed rollout.
The pilot is a test, not an assumption of success. Kentz’s measured view that Jordan isn’t ready for rollout across the district is the kind of caution the paper wants. The paper’s position against federal licensing of AI and broad rules that favor incumbents aligns with Dayton’s approach of evaluating before scaling. The paper’s opposition to letting big companies hide their failures is reflected in the disclosure that chat logs can be viewed by school employees, an external evaluator and SchoolAI staff.
The lack of metrics across the field is the core problem, and Dayton’s pilot is still working through it. The paper wants districts to measure what they are doing and demand transparency from vendors before committing to purchases. Readers should watch for whether Dayton develops clearer metrics for what success looks like, and whether the evaluation arrangement with Kentz remains free from conflict of interest.
Source material: “Welcome to the 'Wild West' of AI in schools: Research is scarce, but experiments abound,” NPR.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

