Improving Human Improvement with AI

In 1962, computing pioneer Douglas Engelbart published his framework, Augmenting Human Intellect. He formulated bootstrapping: systematically improving how we improve.

Engelbart observed that standard tools produce linear gains by automating isolated tasks. Investing in tools that elevate human learning compounds across every discipline. When cognitive acquisition accelerates, downstream breakthroughs in energy, materials, and computing multiply rapidly.

For six decades, educational institutions faced an insurmountable economic constraint. In 1984, educational psychologist Benjamin Bloom identified Bloom’s 2-Sigma Problem. Bloom demonstrated that students tutored one-on-one scored two standard deviations above conventional lecture classrooms. Personal tutoring shifted an average student to the 98th percentile.

Students and researchers in a sunlit modern laboratory collaborating with holographic mathematical models and cognitive AI tutors with dark aurora accents.

The pedagogical solution was proven. Yet the economic reality was prohibitive: providing every student on Earth with an expert human tutor would require tens of trillions of dollars in annual labor.

In September 2026, three convergent milestones demonstrated that this economic barrier has broken:

  • Empirical Parity: The StudentBench benchmark (arXiv:2609.28470), analyzed in Curtis Northcutt’s technical breakdown on X, proved that AI tutors match expert human tutors in measurable GRE learning gains (p = 0.015).
  • Economic Collapse: StudentBench demonstrated a 918 x cost reduction on Gemma 4 31B, reducing inference cost per percentage point gained to $0.0052 versus the $4.81 market reference rate.
  • National Validation: In a 171-school pilot in El Salvador, volunteer secondary students achieved PISA for Schools benchmark scores comparable to Germany and Sweden, as highlighted in President Nayib Bukele’s post on X.

We are witnessing the industrialization of human improvement.

Video preview frame for Sal Khan TED talk discussing Bloom's two sigma problem and AI tutoring.
<span class="play-badge" aria-hidden="true">&#9654;</span>
<span class="play-label">Play Video Explainer: Bloom's 2-Sigma Problem &amp; AI (Sal Khan)</span>

1. The Steepest Cost Collapse in Industrial History

To understand why personalized human improvement is suddenly scalable, examine the macroeconomic context. Modern artificial intelligence is not progressing along an ordinary software adoption curve. It represents the steepest cost collapse of any general-purpose technology in documented economic history.

Research compiled by Epoch AI tracks the cost required to achieve a fixed benchmark of frontier model performance. Between 2023 and 2026, the cost of frontier AI performance fell at 47 % per quarter.

Compounded annually, this represents an annual cost collapse exceeding 92 % per year.

Horizontal bar chart comparing annual cost decline rates across historical technology epochs, showing AI performance cost falling at 92 percent per year.
Figure 1: Annual cost decline rates across historical technology epochs. In log-scale deflation, frontier AI performance costs collapse 54 x faster than electricity, 18 x faster than batteries, and 6 x faster than Moore’s Law.

Compare this rate of deflation against previous technological transformations:

  • Electricity (1892–1973): The real cost of electric power fell at roughly 4.6 % per year over 81 years (Nordhaus 1996), compounding at 1.05 x annually.
  • Lithium-ion Batteries (1991–2024): Battery pack costs declined at 13 % per year over three decades, enabling electric transportation.
  • Computing Hardware (1940–2001): Moore’s Law drove microchip compute costs down by approximately 35 % annually.
  • DNA Sequencing (2001–2025): The transition to next-generation sequencing drove costs down by 45 % per year, outpacing Moore’s Law.
  • Frontier AI (2023–2026): Algorithmic efficiency and specialized silicon have driven capability-adjusted costs down by 92 % per year.

In log-scale cost deflation, AI inference costs are collapsing 54 x faster than electricity and 6 x faster than computing hardware. Epoch AI notes that frontier model estimates carry sensitivity between 42.9 % and 58.0 % quarterly depending on sample weights. In contrast, researchers at MIT FutureTech (Gundlach et al., arXiv:2511.23455) estimate model deflation at a conservative 3 x to 10 x annually. Even at that lower bound, AI capabilities decline faster than any historical manufacturing input.

When applied to instruction, this dynamic converts one-on-one cognitive mentorship from an expensive luxury good into an ambient public utility.

2. StudentBench: Proving Tutoring Parity and the 918x Multiplier

Skeptics long argued that AI models could generate textbook explanations, but could not duplicate the nuanced Socratic feedback of an expert human tutor.

On September 24, 2026, Cleanlab CEO Curtis Northcutt and researchers funded by Handshake AI published StudentBench: AI and human tutoring yield equivalent GRE learning gains (arXiv:2609.28470). Northcutt summarized the results in an in-depth analysis thread on X.

+-------------------------------------------------------------------------------+
| METRIC                       | EXPERT HUMAN TUTOR    | FRONTIER AI TUTOR      |
+------------------------------+-----------------------+------------------------+
| Student Sample Size          | 140 students          | 2,139 students         |
| Unassisted Control Cohort    | 190 students          | (Baseline reference)   |
| Tutoring Interaction Mode    | Live video sessions   | 175,000+ text messages |
| Pooled Learning Gain         | Statistically Equal   | Statistically Equal    |
| Statistical Significance     | p = 0.015 (pooled)    | p = 0.015 (pooled)     |
| Domain Breakdown (7 Areas)   | Won Verbal Section    | Won 5 of 7 categories  |
| Cost per % Point Gained      | $4.81 (market rate)   | $0.0052 (Gemma 4 31B)  |
| Cost Advantage Multiplier    | Baseline (1.0x)       | 917.8x cheaper         |
+-------------------------------------------------------------------------------+

The scale of StudentBench separates it from previous classroom evaluations. The trial encompassed 2,383 total participants across three experimental arms: 2,139 students assigned to AI tutors, 140 students receiving live video coaching from verified human instructors, and 190 unassisted control students.

Over 175,000 text exchanges were recorded in the student-AI arm across Quantitative and Verbal Graduate Record Examination (GRE) sections.

Rigorous Pedagogical Equivalence

The study paired students with either professional human GRE tutors or AI tutoring agents across 13 distinct language models. Both groups completed standardized pre-tests and post-tests to quantify learning gains directly.

The findings verified that student learning gains under AI instruction were statistically indistinguishable from expert human tutoring (p = 0.015 pooled). Human tutors maintained a slight performance advantage on the complex Verbal section, where equivalence fell short of significance (p = 0.085). However, on Quantitative reasoning and across five of the seven evaluated subdomains, the top AI tutors achieved higher average score gains than professional human instructors.

AI tutors do not suffer from fatigue, frustration, or divided attention. They adapt their pacing to the exact student knowledge frontier, correcting conceptual misconceptions without judgment.

The 918x Cost Disruption

The pedagogical parity measured by StudentBench (arXiv:2609.28470) was striking, but the financial disparity documented in Curtis Northcutt’s thread on X was historic:

  • Human Tutoring: Cost $4.81 ($4.80508) per percentage point gained, based on standard market reference rate cards ($75 / hr, Jantzi 2018).
  • AI Tutoring: Cost $0.0052 ($0.00524) per percentage point gained using open-weight Gemma 4 31B inference.

That represents a 917.8 x (rounded to 918 x) reduction in the cost of cognitive capability acquisition as detailed in the StudentBench release.

To put a 918 x multiplier in perspective: an educational intervention that previously cost a school district $100,000 now costs $109. An intensive $5,000 test-preparation curriculum becomes accessible for $5.45.

StudentBench does not claim to have fully resolved Bloom’s 2-Sigma challenge across all grade school subjects. However, it establishes a decisive proof-of-concept: algorithmic instruction can deliver expert-level learning gains while destroying the economic bottleneck that once made personal tutoring impossible.

3. National Proof: El Salvador’s PISA Benchmark

Lab benchmarks provide rigorous control, but do these gains transfer to under-resourced public school systems?

El Salvador explored this question through an ambitious educational initiative. On December 11, 2025, the Salvadoran government announced a formal partnership with xAI to pilot Grok-powered interactive tutors across secondary schools.

The pilot deployed structured AI tutoring to 1,198 volunteer students across 171 public secondary schools. The intervention combined digital tutoring with updated curriculum materials, student laptops, classroom internet access, and teacher coaching.

In June 2026, the OECD administered the standardized PISA for Schools assessment to evaluate the volunteer cohort against international baselines.

The benchmark yielded noteworthy results:

  • Students in the pilot cohort achieved average scores comparable to the 2022 PISA national averages of Germany and Sweden in reading and mathematics.
  • Participating students in underfunded rural schools narrowed historical performance divides through daily Socratic dialogues.
  • While the volunteer cohort benefited from bundled technology and coaching investments, the pilot showed that interactive AI tutors can successfully engage students in low-resource environments.

In September 2026, President Nayib Bukele addressed the results on X, noting that ordinary public school children can achieve world-class results under proper conditions. Bukele stressed that upcoming national assessments will verify whether these gains spread across the entire public system.

Following these initial findings, El Salvador is modernizing public education through the AprendES program. Supported by $501.253 M in World Bank financing, AprendES is upgrading infrastructure across 3,200 prioritized schools serving over 70 % of the nation’s public student enrollment.

When a developing nation can elevate public school performance toward parity with Northern Europe at negligible marginal software cost, educational inequality ceases to be an inevitable condition.

4. Engelbart’s Bootstrapping Loop in Practice

Why does this matter beyond test scores and school curricula?

Return to Douglas Engelbart’s concept of bootstrapping. When you deploy AI to write corporate marketing copy, you gain marginal efficiency. When you deploy AI to tutor human beings, you create an accelerating feedback loop.

       +-------------------------------------------------------+
       |                                                       |
       v                                                       |
+---------------------+       +-----------------------+        |
| 92% Annual Compute  | ----> | 918x Cheaper Socratic |        |
| Cost Deflation      |       | Tutoring Engines      |        |
+---------------------+       +-----------------------+        |
                                          |                    |
                                          v                    |
+---------------------+       +-----------------------+        |
| Accelerated Civil,  | <---- | Broadened Population  |        |
| Scientific & Energy |       | Cognitive Capability  |        |
| Discoveries         |       +-----------------------+        |
+---------------------+                   |                    |
       |                                  |                    |
       +----------------------------------+--------------------+

Consider the compound sequence:

  1. Inference Deflation: Hardware and algorithmic breakthroughs collapse the cost of reasoning tokens by 47 % per quarter.
  2. Accessible Mastery: Personalized 2-Sigma tutoring becomes accessible to every person with an inexpensive smartphone or tablet.
  3. Cognitive Elevation: Millions of young minds develop advanced proficiencies in mathematics, software development, physics, and biological engineering.
  4. Talent Amplification: The global pool of capable researchers, systems engineers, and entrepreneurs expands by orders of magnitude.
  5. Compounded Discovery: These newly empowered thinkers invent better energy infrastructure, superior agricultural techniques, and more efficient computing architectures.

This is the definition of bootstrapping: using technology to improve human capability, so that humans can improve technology faster.

5. Strategic Implications for Engineers and Builders

For developers, educators, and technology leaders, the convergence of StudentBench, Epoch AI metrics, and El Salvador’s pilot dictates three strategic imperatives:

1. Shift Focus from Automation to Amplification

Replacing human workers generates one-time margin gains for software companies. Augmenting human cognition expands the entire macroeconomic frontier. Build interfaces that coach, challenge, and elevate human operators rather than treating them as passive consumers.

2. Design for Socratic Feedback, Not Passive Answers

StudentBench succeeded because the AI agents functioned as tutors rather than cheat engines. They prompted students with targeted questions, identified flawed assumptions, and forced active recall. Systems that give answers degrade human intellect; systems that ask questions expand it.

3. Exploit the Inference Cost Curve

Do not design educational architectures around today’s API prices. With performance costs declining 92 % each year, computational constraints that seem expensive today will become trivial within eighteen months. Plan systems that provide deep multi-turn reasoning and iterative cognitive scaffolding.

Conclusion: The Ultimate Dividend

Throughout history, civil progress was bottlenecked by human cognitive bandwidth. Only a tiny fraction of the human population enjoyed the leisure and resources to master advanced mathematics, literature, and physical sciences.

Bloom’s 2-Sigma Problem proved that human potential is vast, but society lacked the economic capacity to nurture it individually.

By combining the steepest cost collapse in industrial history with empirically validated Socratic agents, we have unlocked Engelbart’s bootstrapping dream. We are no longer merely building smarter machines.

We are improving human improvement itself.

Back to Articles Back to homepage