Chapter 11: Assessment Beyond the Bubble Sheet
Measuring What Matters in the Fourth Industrial Revolution
The superintendent leaned back in his chair and said it as plainly as anyone has ever said anything about the most contentious topic in American public education: "Assessment is multifaceted" (Martin, 2022, p. 96).
I waited for the caveat. The hedge. The obligatory acknowledgment that standardized tests are imperfect but necessary, that we must work within the system, that change takes time. Every superintendent I have ever spoken to about assessment eventually arrives at the hedge. It is the professional survival instinct of a leader who knows that their job security, their district's reputation, and their schools' funding are all tethered to a number on a state report card.
He did not hedge.
Instead, he described a system I had never seen in a California public school district -- not because it was radical, but because it was coherent. At Innovation USD, standardized test scores were accepted as a reality of the accountability landscape. Nobody pretended they did not exist. Nobody mounted a quixotic campaign to abolish them. But neither did anyone confuse them with evidence of learning. The scores were one data point in a constellation that included project showcases where students presented original research to authentic audiences, growth portfolios that tracked each learner's development across multiple dimensions over time, goal-oriented feedback sessions where students set their own learning targets and reflected on their progress, and community impact presentations where young people demonstrated what they had built, designed, or solved for their neighborhoods (Martin, 2022).
A third-grader presented her water conservation project to the city council. A high school student's portfolio included a business plan he had actually executed. And the district could predict outcomes -- not just measure them after the fact. "How do we know that we've arrived?" the superintendent asked me. "What I say through this process is, if we can be predictive of outcomes is when I know. I want to be predictive of the outcomes of how our students are going to do based on this work here" (Martin, 2022, p. 97).
This was not assessment reform. This was a fundamentally different theory of what evidence means, who produces it, and what it is for. And it emerged not from a testing vendor's sales pitch or a state mandate or a thought leader's keynote -- it emerged from a superintendent who understood that you cannot transform pedagogy with an Education 2.0 measurement system and call it progress. If we are serious about preparing students for the Fourth Industrial Revolution, then we must be equally serious about measuring what that preparation actually looks like. This chapter is about how to do that -- honestly, practically, and without pretending that the accountability landscape does not exist.
The Eugenic Architecture of the Bubble Sheet
Before we can build a better assessment system, we must understand what we are building it against. The standardized testing regime that dominates American public education was not designed to measure learning. It was designed to sort human beings.
The SAT -- the test that has shaped college admissions, school rankings, and public perceptions of "quality education" for nearly a century -- was created by Carl Brigham, an avowed eugenicist. Brigham was a disciple of Robert Yerkes, who administered the Army Alpha and Beta intelligence tests during World War I to classify soldiers by cognitive ability. The tests were explicitly designed to demonstrate the intellectual superiority of Northern European men and the inferiority of immigrants, Black Americans, and anyone else who did not conform to the Anglo-Saxon norm (Brigham, 1923). When Brigham adapted the Army tests into the SAT for the College Entrance Examination Board in 1926, he did not abandon the eugenicist framework. He refined it. The test measured a particular kind of cognitive performance -- the kind cultivated by affluent white households with access to books, leisure, and formal language patterns -- and called it "aptitude" (Au, 2009).
This is not ancient history. It is the architecture of the system we are still using. The SAT has been revised, renamed, and restructured multiple times, and the College Board has formally distanced itself from Brigham's legacy. But the fundamental logic of standardized testing -- that human ability can be captured by a single score, that the score predicts future performance, and that the score should determine access to opportunity -- is Brigham's logic. It is the logic of enclosure: a mechanism for sorting students into categories that predict and reproduce existing hierarchies of race and class (Au, 2009; Kendi, 2016).
The No Child Left Behind Act of 2001 weaponized this logic at scale. By tying school funding, principal tenure, and district reputation to standardized test performance, NCLB made the bubble sheet the only measure that mattered. Schools that served predominantly Black and Latino students -- schools that had been systematically deprived of resources for generations -- were now punished for the outcomes that deprivation produced. The testing regime did not reveal inequity. It enforced it. Race to the Top, which followed under the Obama administration, did not dismantle this structure; it intensified it by adding teacher evaluation to the testing stakes (Ravitch, 2013; Darling-Hammond, 2015).
Here is the paradox that every superintendent in America must now confront: inquiry-based learning -- the very pedagogy that develops the critical thinking, creativity, collaboration, and complex problem-solving that the Fourth Industrial Revolution demands -- does not improve standardized test scores. Kingston's (2018) comprehensive review of project-based learning research found that the relationship between PBL and standardized test performance was, in his word, "unsubstantiated." Students who learn through inquiry develop deeper understanding, stronger transfer skills, and more durable knowledge -- but they do not fill in bubbles faster. The assessment system punishes the pedagogy that prepares students for the future. It rewards the pedagogy that prepares them for 1926.
This is not an inconvenient finding. It is the central structural contradiction of American public education in the AI era. And no superintendent can transform their district without confronting it directly.
Innovation USD's Multifaceted Assessment Model
What I observed at Innovation USD was not a district that had abandoned accountability. It was a district that had expanded its definition of evidence -- deliberately, strategically, and in a way that honored the state accountability framework while refusing to be imprisoned by it.
The model operated on four interconnected levels.
Level One: Standardized Testing as Baseline, Not Ceiling. The superintendent never dismissed state test data. He used it. But he used it the way a physician uses a blood pressure reading -- as one vital sign among many, meaningful only in combination with other indicators, and dangerous when treated as the sole diagnosis. At Innovation USD, CAASPP scores were disaggregated by student group, analyzed for trends, and incorporated into school planning. But they were explicitly positioned as a lagging indicator -- a measurement of where students had been, not where they were going. "I'm not saying to be perfect," the superintendent told his principals. "I'm just saying go through the process so we can learn" (Martin, 2022, p. 97).
Level Two: Project Showcases and Performance-Based Assessment. Every school in the district held regular student showcases where learners presented original work to authentic audiences -- parents, community members, city officials, industry partners. These were not science fairs with tri-fold posterboards. They were demonstrations of sustained inquiry: research questions developed by students, investigations conducted over weeks or months, findings presented with evidence, and solutions proposed for real community problems. The showcases served a dual function. They provided assessment data -- teachers and external evaluators used common rubrics to evaluate student work against competency standards. And they built community buy-in, because when a parent watches their child explain a water quality analysis to a city council member, the question "But what were the test scores?" becomes less compelling.
Level Three: Growth Portfolios and Goal-Oriented Feedback. Each student maintained a portfolio that tracked growth across multiple dimensions -- not just academic content but collaboration skills, communication, creative thinking, and self-direction. These portfolios were not scrapbooks. They were structured reflective documents where students identified their own learning goals, collected evidence of progress toward those goals, and engaged in regular conferences with teachers to assess their trajectory. The shift here was profound: students were not being assessed. They were assessing themselves, with teacher guidance. Assessment was not something done to them. It was something they owned.
Level Four: Quarterly SPSA Reviews and Predictive Analytics. Perhaps the most distinctive feature of Innovation USD's assessment model was the quarterly Single Plan for Student Achievement review. Every principal sat down with the superintendent, disaggregated data in hand, and reflected on the progress of each student segment -- not just overall averages, but specifically Latino students, African American students, English learners, students with disabilities. "You must name it," the superintendent had insisted from the beginning. "When it's not named, and everything is meshed together, then you don't have a specific game plan" (Martin, 2022, p. 79). The SPSA reviews were not compliance exercises. They were accountability conversations rooted in equity -- conversations where a principal had to explain, with evidence, what was working and what was not working for the students who had historically been least well served.
What made this model work was not any single component. It was the coherence between them. The test data informed the SPSA reviews. The SPSA reviews informed the professional development. The professional development shaped the pedagogy. The pedagogy produced the student work in the portfolios and showcases. And the portfolios and showcases generated the evidence that the test data alone could never capture. It was a system -- not a collection of add-ons layered over a testing regime.
Rounds of Inquiry: Seeing the Change You Cannot Test
One of the most powerful assessment practices I encountered at Innovation USD had nothing to do with student performance data. It was called "rounds of inquiry," and it measured something that no standardized test can capture: whether the adults in the building had changed.
Here is the premise. A district can adopt inquiry-based learning as a policy. It can purchase curriculum materials, restructure schedules, and send teachers to PBL workshops. But none of that matters if the daily interactions between teachers and students remain unchanged. The question that rounds of inquiry were designed to answer was deceptively simple: When I walk into a classroom, what do I see the teacher doing, and what do I see the students doing?
District leaders conducted regular equity visits -- structured classroom observations focused not on evaluating individual teachers but on gathering systemwide evidence of pedagogical shift. The observation protocols were specific. Observers noted the ratio of teacher talk to student talk. They documented whether the learning task required recall or construction. They recorded whether students were working individually in silence or collaborating in groups. They observed whether the content connected to students' cultural backgrounds and community contexts. And critically, they looked for evidence that historically marginalized students -- the students named in the district's theory of action -- were actively engaged as thinkers and contributors, not passive recipients of instruction (Martin, 2022).
"Evidence of impact measured at multiple levels: lesson delivery change, student response, progress monitoring," the Director of Instruction and Innovation explained, describing what the rounds revealed (Martin, 2022, p. 112). This was a three-layer model of evidence. The first layer was teacher behavior: Had the teacher shifted from lecturer to facilitator? Was the classroom organized for inquiry? Were the tasks open-ended? The second layer was student response: Were students asking questions or only answering them? Were they engaged in sustained investigation or compliant seat-time? Were the students who had historically been silenced now speaking? The third layer was progress monitoring: Over time, across multiple rounds, was the pattern shifting? Were classrooms moving toward inquiry, or had the professional development evaporated the moment the workshop ended?
This three-layer approach represents a fundamentally different theory of assessment -- one that measures the conditions for learning, not just the outputs of learning. It recognizes that student outcomes are downstream of adult behavior, and that measuring adult behavior change is both more actionable and more honest than waiting for a test score to move. A superintendent who sees, through rounds of inquiry, that 60% of classrooms in a school still operate as lecture-and-worksheet environments knows exactly where to direct coaching resources. A superintendent who sees that students of color are disproportionately assigned to passive roles in "collaborative" activities knows exactly which equity conversation needs to happen next. The data is immediate, specific, and connected to action.
Rounds of inquiry also served a cultural function. When district leaders showed up in classrooms not to evaluate but to learn alongside teachers, it modeled the very inquiry stance the district was asking teachers to adopt with students. The superintendent was explicit about this: "Remember the whole design of the middle? That's brought me more credibility than anything else, moving us farther than anything else" (Martin, 2022, p. 88). Leading from the middle meant that assessment -- of teachers, of students, of the system itself -- was a shared enterprise, not a top-down judgment.
Fullan and Quinn (2016) call this "accountability from the inside out" -- a model where professionals hold themselves and each other accountable through transparent reflection rather than external surveillance. It is the difference between a compliance audit and a learning walk. Both involve observation. Only one produces change.
AI-Enhanced Assessment: Promise, Peril, and the Centering of Student Voice
If we accept the premise of the previous chapter -- that pedagogy must precede technology -- then the question for assessment in the AI era is not "How can AI improve testing?" but rather "How can AI enhance the multifaceted assessment model that inquiry-based learning demands?"
The answer is both promising and perilous, and the difference between the two depends entirely on who controls the technology and whose voice it centers.
The promise is real. AI-powered adaptive assessment systems can provide formative feedback in real time -- not end-of-year summative judgments but moment-by-moment guidance that helps students adjust their learning while the learning is still happening. A meta-analysis of 51 studies published in Humanities and Social Sciences Communications found that ChatGPT had a large positive impact on learning performance (g = 0.867) and moderate positive impacts on learning perception and higher-order thinking (Nature, 2025). A Harvard randomized controlled trial found that students using an AI tutor learned significantly more in less time compared to traditional active learning and reported higher engagement and motivation (Kestin et al., 2025). Google DeepMind's study of 165 students across five UK secondary schools found that students using supervised AI tutoring solved new problems at a higher rate than those with human tutors (Google DeepMind, 2025).
These findings suggest that AI, when deployed within an inquiry-based pedagogical framework, can function as what Vygotsky (1978) described as a more knowledgeable other -- a responsive partner that scaffolds learning at the edge of a student's current capability. AI can analyze a student's growth portfolio and identify patterns invisible to the human eye. It can provide differentiated feedback on a project-based learning task at a scale that no single teacher can manage across 35 students. It can generate formative assessment prompts that adapt in real time to a student's evolving understanding.
But the peril is equally real. When AI is deployed without an inquiry-based pedagogy -- which is to say, in most American classrooms right now -- it does not enhance assessment. It automates the deficit model. RAND Corporation (2025) found that teachers used AI most frequently for making worksheets and modifying materials, not for facilitating inquiry or providing formative feedback on complex student work. The Brookings Institution's landmark 2025 report, based on a year-long global study across more than 50 countries, concluded that "the risks of utilizing generative AI in children's education overshadow its benefits" given current patterns of use -- patterns where AI "replaces thinking instead of extending it" (Brookings, 2025, p. 4). A 2024 study found that high school math students tutored by ChatGPT initially scored better but the benefit soon evaporated because they had failed to acquire conceptual understanding (Education Week, 2025). And a 2025 experiment discovered a clear "cognitive cost" to receiving AI help with writing essays, with a separate study finding that more AI use was "associated with lower critical thinking skills" (Education Week, 2025).
The pattern is unmistakable, and it mirrors the pedagogy-first principle from Chapter 10: AI in assessment amplifies whatever assessment philosophy it encounters. If the philosophy is sorting and deficit, the AI sorts and deficits faster. If the philosophy is growth and inquiry, the AI supports growth and inquiry at scale.
The centering of student voice is the differentiator. At Innovation USD, the superintendent's assessment model worked because it treated students as agents of their own learning, not objects to be measured. Growth portfolios required students to set goals and reflect on progress. Project showcases required students to present and defend their work. The quarterly SPSA reviews included data about student experience, not just student performance. This is what the OECD Learning Compass 2030 framework calls "student agency" -- the capacity of learners to set goals, reflect on their actions, and act responsibly to effect change (OECD, 2019).
AI can amplify student agency or extinguish it. An AI system that generates a score and a remediation pathway -- without any student input into the goals, the process, or the meaning of the data -- is not assessment. It is surveillance. An AI system that helps a student track their own growth, identify their own strengths, and articulate their own next steps is not automation. It is empowerment. The technology is the same. The philosophy makes the difference.
Khan Academy's Khanmigo, which expanded from 68,000 users to over 700,000 between 2023-24 and 2024-25 (Khan Academy, 2025), represents an early attempt to build AI tutoring that centers student engagement rather than content delivery. But even Khanmigo's developers acknowledge the challenge: the biggest barrier to meaningful AI tutoring is not the technology but student engagement, with some students responding to the AI tutor with "I don't know" or similar non-engaged responses (Michigan Virtual, 2025). This is not a technology problem. It is a pedagogy problem. Students who have spent their entire school careers in banking-model classrooms do not suddenly become inquiry-driven agents because an AI asks them a question. The pedagogical foundation must be built first. Then the AI can extend it.
The Accountability Tightrope: Working Within the System While Building Beyond It
Every superintendent reading this chapter is thinking the same thing: "This sounds wonderful. But my board wants test scores. My state wants test scores. My community compares our scores to the district next door. I cannot ignore the accountability system, no matter how flawed it is."
I know. And I am not asking you to.
The superintendent at Innovation USD did not ignore test scores. He used them. He disaggregated them by student group with more precision than most of his peers. He held principals accountable for growth among specific populations. He presented testing data to his school board quarterly. He played the accountability game, and he played it well -- Innovation USD was one of the top-performing districts in California, with accountability measures 39% above state averages (Martin, 2022).
But he did something else, simultaneously. He built a parallel evidence system -- one that captured what the tests could not -- and he systematically taught his board, his community, and his staff to value that evidence alongside the test data. This is the strategic insight that most superintendents miss: you do not have to choose between accountability and authenticity. You have to build them both, and you have to narrate the relationship between them.
Here is how that works in practice.
First, present the test data honestly and disaggregated. Do not hide it, spin it, or dismiss it. Board members and community members can tell when a superintendent is deflecting. Instead, own the data, name the gaps by student group (as the Innovation USD superintendent insisted), and frame the scores as one measure of progress -- necessary but insufficient. This builds credibility. You cannot ask people to value alternative evidence if you have not first demonstrated that you take the traditional evidence seriously.
Second, introduce complementary evidence in every board presentation and community communication. When you report CAASPP scores, also report the number of students who presented original work to external audiences. Report the percentage of classrooms where rounds of inquiry found student-centered pedagogy in practice. Report portfolio completion rates and the percentage of students who met self-set learning goals. Report community impact metrics: projects that solved real problems, partnerships that produced real outcomes. Present this evidence with the same rigor, the same disaggregation, and the same visual clarity as the test data. Over time, the complementary evidence becomes expected, not supplementary.
Third, use student voice as the most compelling evidence. When a ninth-grader stands before the school board and explains the civic research project she designed, conducted, and presented to city officials -- without notes, with evidence, with passion -- the board does not ask about her CAASPP score. They do not need to. The evidence of learning is standing in front of them. Innovation USD's parent affinity groups -- the African American Parent Advisory Committee, the Latino Parent Alliance -- became powerful allies in this narrative. When parents saw their children engaged as thinkers and creators, they became advocates for the multifaceted model. "Sell, don't tell," the superintendent advised. Take people to see the work. Let the evidence speak for itself (Martin, 2022).
Fourth, align the alternative assessment system to recognized frameworks. The World Economic Forum's Education 4.0 Taxonomy (2023) provides a comprehensive competency structure that complements state content standards. The OECD Learning Compass 2030 offers an internationally recognized framework for student agency, well-being, and competencies (OECD, 2019). The OECD and European Commission's draft AI Literacy Framework (2025) defines domains -- engaging with AI, creating with AI, managing AI, designing AI -- that represent measurable competencies no standardized test captures. When your complementary assessment system is aligned to these frameworks, you are not presenting "soft" alternatives to "hard" data. You are presenting internationally recognized competency measures alongside state standardized measures. Both are rigorous. Both are necessary. Neither is sufficient alone.
Fifth, build the narrative over time. Assessment transformation is not a single board presentation. It is a multi-year campaign to shift what a community considers evidence of a good education. The superintendent at Innovation USD understood that the community had passed two bond measures at 72% approval -- nearly $700 million -- not because the test scores were strong (though they were) but because the community had been shown what learning looked like when it was real. The "sell, don't tell" approach -- taking community members, board members, and skeptics to observe inquiry-based classrooms, to attend student showcases, to sit in on portfolio conferences -- created a constituency for authentic assessment that no test score could produce.
This is not easy work. It requires a superintendent who is willing to be evaluated on more than one metric and courageous enough to ask their board to do the same. It requires patience, because the parallel evidence system takes years to build credibility. And it requires the kind of moral clarity that Chapter 6 described -- the willingness to name what is being measured, for whom, and why. But it is the only path that honors both the reality of the accountability landscape and the reality of what students actually need.
Competency-Based Assessment for the AI Era: What Should We Actually Measure?
If we are building an assessment system for the Fourth Industrial Revolution, we must answer a question that the current system never asks: What, specifically, should we be measuring?
The World Economic Forum's (2025) Future of Jobs Report provides one answer: 39% of workers' key skills are expected to change by 2030, and skills in AI-exposed jobs are changing 66% faster than in less-exposed occupations. The most in-demand competencies -- analytical thinking, creative thinking, resilience, flexibility, technological literacy, leadership, collaboration -- are precisely the competencies that standardized tests do not and cannot measure. You cannot bubble in "resilience." You cannot multiple-choice "collaboration." You cannot time-limited-essay "creative thinking" in any authentic way.
The WEF's Education 4.0 Taxonomy (2023) offers a structured framework for competency-based assessment, organizing aptitudes into a tree structure that facilitates dialogue between educators, policymakers, and employers. The OECD Learning Compass 2030 adds student agency, well-being, and the capacity to navigate complexity and ambiguity as core competencies (OECD, 2019). Together, these frameworks define what a Fourth Industrial assessment system should measure:
None of these competencies is mysterious. None is unmeasurable. Each can be assessed with rubrics, portfolios, performance tasks, and structured reflections -- assessment methods that already exist, that are already validated, and that are already in use in districts like Innovation USD. The obstacle is not methodological. The obstacle is political: a testing infrastructure that generates billions of dollars in revenue for testing companies, that provides a simple narrative for politicians and media, and that reduces the magnificent complexity of human learning to a single number on a page.
Competency-based assessment does not produce single numbers. It produces stories -- richly documented, evidence-backed stories of what a student can do, how they think, how they collaborate, and how they grow. For some stakeholders, this complexity is threatening. For the students and communities who have been sorted and discarded by the bubble sheet for a century, it is liberation.
Discussion Questions
Practitioner Tool: The Multifaceted Assessment Design Guide
Purpose: A framework for building a complementary assessment system alongside existing state accountability measures -- not replacing mandated tests, but ensuring they are contextualized within a broader evidence ecosystem.
Component 1: Portfolio Assessment Rubric (Aligned to 4IR Competencies)
| Competency Domain | Emerging (1) | Developing (2) | Proficient (3) | Distinguished (4) |
|---|---|---|---|---|
| Inquiry & Analytical Thinking | Identifies a topic; relies on single source | Develops a question; gathers multiple sources | Designs investigation; evaluates evidence critically | Original research question; synthesizes across sources; proposes evidence-based solutions |
| Creative & Design Thinking | Reproduces existing solutions | Modifies existing solutions for new context | Generates novel approaches; iterates based on feedback | Designs innovative solutions; documents iterative process; transfers across domains |
| Collaboration & Communication | Participates when directed | Contributes ideas; listens to peers | Facilitates group process; communicates to varied audiences | Navigates conflict productively; adapts communication to audience; amplifies others' contributions |
| AI Literacy & Tech Fluency | Uses AI tools as directed | Evaluates AI outputs for accuracy | Critically assesses AI bias and limitations; creates with AI | Designs AI-integrated solutions; advocates for responsible use; teaches others |
| Self-Direction & Agency | Completes assigned tasks | Sets goals with teacher support | Sets independent goals; monitors progress; adjusts approach | Designs own learning pathway; mentors peers; reflects with sophistication |
| Cultural Competence & Civic Engagement | Acknowledges different perspectives exist | Seeks multiple perspectives on issues | Integrates diverse perspectives into work; engages community | Leads community impact projects; advocates for equity; connects learning to civic action |
Component 2: Project Showcase Protocol
- Frequency: Minimum once per semester, with school-level showcases quarterly
- Audience: Must include at least one external stakeholder group (parents, community members, industry partners, municipal officials)
- Student role: Students present, explain, and defend their work; respond to audience questions
- Evaluation: Common rubric used by teacher + at least one external evaluator
- Documentation: Video or written record entered into student portfolio
Component 3: Student Self-Assessment Instrument
- Administered quarterly alongside portfolio conferences
- Students identify: (a) learning goals set at start of period, (b) evidence of progress toward each goal, (c) areas of growth, (d) revised goals for next period
- Teacher role: facilitator and thought partner, not evaluator
- Data aggregated at school and district level to track student agency development
Component 4: Community Impact Metrics
- Number of student projects addressing real community needs
- Number of external partnerships (civic, industry, nonprofit) engaged in student learning
- Documented community outcomes from student work (policy changes, environmental improvements, services delivered)
- Student reflections on civic learning and community connection
Component 5: Data Integration Model
All five data streams -- standardized test scores, portfolio assessments, showcase evaluations, student self-assessments, and community impact metrics -- are reported together in a unified student profile. No single measure stands alone. The profile tells a story: who this student is, what they can do, how they think, how they have grown, and what they are ready for next.
References
Au, W. (2009). Unequal by design: High-stakes testing and the standardization of inequality. Routledge.
Brigham, C. C. (1923). A study of American intelligence. Princeton University Press.
Brookings Institution. (2025). A new direction for students in an AI world: Prosper, prepare, protect. https://www.brookings.edu/articles/a-new-direction-for-students-in-an-ai-world/
Center for Democracy & Technology. (2025). Hand in hand: Schools' embrace of AI connected to increased risks to students. https://cdt.org/insights/hand-in-hand-schools-embrace-of-ai-connected-to-increased-risks-to-students/
Darling-Hammond, L. (2015). The flat world and education: How America's commitment to equity will determine our future. Teachers College Press.
Education Week. (2025, September). AI in the classroom is often harmful. Why are educators falling prey to the hype? https://www.edweek.org/technology/opinion-ai-in-the-classroom-is-often-harmful-why-are-educators-falling-prey-to-the-hype/2025/09
Freire, P. (2000). Pedagogy of the oppressed (30th anniversary ed.). Continuum. (Original work published 1970)
Fullan, M., & Quinn, J. (2016). Coherence: The right drivers in action for schools, districts, and systems. Corwin.
Gallup & Walton Family Foundation. (2025). Three in 10 teachers use AI weekly, saving six weeks a year. https://news.gallup.com/poll/691967/three-teachers-weekly-saving-six-weeks-year.aspx
Google DeepMind. (2025). AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms. https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_nov25.pdf
Hattie, J. (2012). Visible learning for teachers: Maximizing impact on learning. Routledge.
Kendi, I. X. (2016). Stamped from the beginning: The definitive history of racist ideas in America. Nation Books.
Kestin, G., Miller, K., McCarty, L. S., Callaghan, K., & Deslauriers, L. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. https://www.nature.com/articles/s41598-025-97652-6
Khan Academy. (2025). Khan Academy annual report: SY24-25. https://annualreport.khanacademy.org/
Kingston, S. (2018). Project based learning and student achievement: What does the research tell us? PBL Works. https://www.pblworks.org/research/project-based-learning-and-student-achievement
Martin, M. G. (2022). The Fourth Industrial Superintendent [Doctoral dissertation, University of Southern California]. USC Digital Library.
McKinsey Global Institute. (2023). The economic potential of generative AI: The next productivity frontier. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier
Mehta, J., & Fine, S. M. (2019). In search of deeper learning: The quest to remake the American high school. Harvard University Press.
Michigan Virtual. (2025). Have you considered AI in your classroom? A Khanmigo pilot story. https://michiganvirtual.org/blog/have-you-considered-ai-in-your-classroom-a-khanmigo-pilot-story/
Nature. (2025). The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: Insights from a meta-analysis. Humanities and Social Sciences Communications. https://www.nature.com/articles/s41599-025-04787-y
OECD. (2019). OECD Learning Compass 2030. https://www.oecd.org/en/data/tools/oecd-learning-compass-2030.html
OECD & European Commission. (2025). Empowering learners for the age of AI: An AI literacy framework for primary and secondary education [Review draft]. https://ailiteracyframework.org/wp-content/uploads/2025/05/AILitFramework_ReviewDraft.pdf
Ravitch, D. (2013). Reign of error: The hoax of the privatization movement and the danger to America's public schools. Knopf.
RAND Corporation. (2025). AI use in schools is quickly increasing but guidance lags behind. https://www.rand.org/pubs/research_reports/RRA4180-1.html
Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes. Harvard University Press.
World Economic Forum. (2023). Defining Education 4.0: A taxonomy for the future of learning. https://www3.weforum.org/docs/WEF_Defining_Education_4.0_2023.pdf
World Economic Forum. (2025). The Future of Jobs Report 2025. https://www.weforum.org/publications/the-future-of-jobs-report-2025/