The AI Revolution in Education with Shawn Jansepar, Director of Engineering at Khan Academy
Summary
Khan Academy’s central bet is that GPT-4 has turned one-to-one AI tutoring from a hand-wavy aspiration into a deployable product, though not yet a proven substitute for an expert human. Sean Jancipar predicts that within 10 years no child will learn without an always-available tutor that knows their learning history and, with permission, their interests. The ambition is Bloom’s two-sigma upside; the hedge is explicit: whether AI can match human-tutor outcomes “still remains to be seen.”
Khanmigo’s first-mover advantage came from privileged model access, institutional trust, and a deadline-driven sprint to GPT-4’s March 14 launch. Khan Academy moved from an October Slackbot encounter with what felt like an “omnipotent being” to a Chrome-extension prototype, student testing, OpenAI red-teaming, and a company-wide January hackathon. Sean said that after the launch, some observers asked why they should compete with a product that had already executed well and had the educational brand and trust.
The model stack remains deliberately GPT-4-heavy because reliable Socratic instruction matters more than premature cost optimization. GPT-3.5 is “basically a no-go” for tutoring because it too readily ignores instructions not to reveal answers, although it handles lower-stakes tasks such as extracting conversation insights and opt-in student interests. The operating principle is to “focus on finding the magic”: better to delight 10 users than ship mediocrity to 10,000, then use cheaper models where quality permits.
Khanmigo’s differentiation sits above the foundation model in educational context, expert-designed behavior, and multi-call reasoning. Exercises supply the question, correct answer, and authored hints; for math, Khanmigo first privately analyzes whether the student is right and why, then feeds that analysis into a separate tutoring response. Khan Academy also helped label roughly 100 pre-release questions, each with many variations, for OpenAI, improving GPT-4’s tutoring continuity rather than its raw arithmetic.
Safety is treated as a product layer rather than a claim that GPT-4 itself is jailbreak-proof. Every message can be wrapped with on-task instructions and passed through OpenAI’s moderation API, while teachers and parents can inspect student logs. Current personalization is modest, but opt-in interests and cross-session learning memory are on the roadmap; the product already warns that “Khanmigo makes mistakes sometimes” and explains why.
Distribution is designed around schools and teacher augmentation, not replacing classrooms with solitary AI use. Khan Academy charges districts per student per month, targets schools with high free-and-reduced-lunch populations, and uses discounts or local corporate sponsors when districts cannot afford access; Nathan obtained individual access through a recurring $9 monthly donation. Marc Bhargava emphasized that the goal is fewer simultaneous raised hands, more time for project-based teaching, native-language help for English learners, and personalized intervention when the AI cannot resolve a problem—not replacing teachers or classrooms.
The investable outcome remains efficacy, and Khan Academy has not yet established it for Khanmigo. Engagement and time spent learning are early signals, but the intended proof is a MAP Growth comparison between students using ordinary Khan Academy and a comparison cohort using Khanmigo alongside it. The roadmap—voice, scanned homework, handwritten-work interpretation, differentiated groups, collaborative stories, and AI-facilitated debates—expands the surface area, but Sean closes with the qualified view that he thinks it can help change the world.
Deep dive
1. GPT-4 turned a distant tutoring thesis into an immediate product decision
In roughly October, before ChatGPT launched, Sean Jancipar was unexpectedly added to a Slack channel containing GPT-4. Thirty minutes with what felt like an “omnipotent being” sent him walking around the block to process the implications; experience later corrected the first impression: “We don’t have AGI,” even if the model still looked world-changing.
Khan Academy’s initial instinct was to keep the technology away from learners. The safer applications seemed to be content production, teacher lesson plans, or other staff tools, because direct student access raised unresolved risks; discovering OpenAI’s moderation API made a guarded learner product seem more plausible.
The breakthrough prototype came from Khan Academy team member Jessica, who created a smarter search or “concierge” concept for directing learners to content. After weekly status meetings went nowhere, the teams switched to virtual rapid prototyping, and a scrappy Chrome extension gave the model each Khan Academy page’s context: a learner could say, “I don’t really get this,” and the “guide on the side” understood what “this” meant.
2. Student demand crystallized around relevance, not merely correct answers
The first student trial took place at Khan Lab School with advance parental consent. Students expected a basic, disappointing chatbot because ChatGPT was not yet public; instead, they were startled by its competence and gravitated most strongly toward one button: “Why should I care about this?”
Jancipar treated that preference as an indictment of conventional instruction: the feature’s popularity “goes to show you how good of a job we do in education” explaining why material matters. He later envisioned using opt-in interests such as soccer to make explanations more personally motivating.
Khan Academy then worked inside OpenAI’s offices on red-teaming, roadmap design, and math reliability. By January, the organization brought much of its staff under NDA for a company-wide hackathon, narrowed the strongest experiments, and committed a large share of engineering, design, and product capacity to launching beside GPT-4 on March 14.
Model selection stayed uncertain until late. OpenAI was producing weekly versions; some partners favored a model stronger at storytelling, while Khan Academy wanted the math-oriented candidate. A weekend test by leaders, engineers, and designers found that the eventual production model offered “the best of both worlds,” while the backup model was less satisfactory to Khan Academy.
3. Mastery learning supplies the causal thesis behind the AI tutor
Jancipar anchored the product in Benjamin Bloom’s two-sigma study, long part of Khan Academy’s pedagogy. He said he thought mastery-based learning alone could deliver at least a one-standard-deviation improvement even with one teacher for 30 students, while one-to-one tutoring combined with mastery represented the ideal result.
Mastery means not advancing until the prerequisite is understood or at least proficient. A student receiving a C in Algebra 1 still enters Algebra 2 under the conventional system, despite carrying material gaps; missing one lecture can similarly make the next segment unintelligible because every later step assumes the lost one.
Khan Academy videos already softened that failure mode: learners can pause, rewind, and retry “with no judgment.” An AI tutor adds immediate diagnosis, conversation, and eventual memory of the learner’s history—the capacity to notice a specific weakness and address it before another layer of curriculum is built on top.
Jancipar’s 10-year prediction is that every learner will have an on-demand tutor that knows their history and, with consent, interests such as the Avengers or soccer. Yet he did not claim the two-sigma result in advance: advanced human experts remain better today, GPT-4 still makes math mistakes, and perfect mathematical reliability remains “a research problem that is not solved yet.”
4. GPT-4 carries tutoring while cheaper models handle peripheral work
Khanmigo is still overwhelmingly powered by GPT-4. Its ability to follow the instruction “be Socratic and do not give away the answer” was essential; GPT-3.5 does not follow that constraint reliably enough and is “basically a no-go” for core tutoring. Llama 2 looked worth exploring, but roadmap opportunity cost had kept it from becoming a priority.
Lower-stakes activities form what Jancipar called a “demo disc”: learners can co-write a story or converse with a historical figure before those interactions become embedded in relevant lessons. Some tolerate GPT-3.5 because strict instructional control matters less than it does beside an exercise, quiz, article, or video.
Cost optimization follows product discovery. Jancipar invoked the startup progression from zero to one and the principle that “10 people absolutely love your product” beats a mediocre product for 10,000; the team’s phrase is, “Let’s just focus on finding the magic.” GPT-3.5 is used where possible for teacher insights, off-topic detection, and opt-in interest extraction because GPT-4 is much more expensive.
The essay tool illustrates the move beyond chat: Khanmigo can highlight passages needing grammar or spelling work rather than describing locations awkwardly in a back-and-forth transcript. Nathan added a model-diversity datapoint from his own workflow—Claude 2 had produced a better stylistic starting point for podcast introductions than GPT-4 when given past samples and the current transcript.
5. Educational experts shaped tutoring behavior before metrics could mature
Jancipar clarified an apparent contradiction: production Khanmigo is not fine-tuned and repeatedly passes article or exercise context, but before GPT-4’s release Khan Academy contributed directly to OpenAI’s reinforcement-learning process. Sal Khan and content specialists evaluated roughly 100 questions, including many variations, distinguishing bad, good, and intermediate tutoring responses.
The recurring defect was conversational, not simply computational. After asking for a next step, the model might see the student’s equation and treat it as an unrelated new problem rather than an answer to its own question. Khan Academy’s feedback materially improved that continuity—making GPT-4 better at tutoring math, Jancipar stressed, not necessarily better at doing arithmetic itself.
Product development likewise shifted from a waterfall handoff toward demo-driven work inspired by Creative Selection, a book about building the iPhone. Designers, engineers, product managers, learning scientists, and content specialists were put together to prototype, demonstrate, critique, and repeat because no complete requirements document could specify the desired tutor in advance.
The management doctrine became “data informed” rather than universally “data driven.” The team still labels conversations to measure when Khanmigo helps or fails, but expert judgment shortens the loop for reversible prototypes; formal efficacy claims remain the exception where rigorous data is indispensable. Nathan’s warning was apt: aggregate benchmarks can mislead unless somebody reads the underlying transcripts.
6. Reversible experiments moved fast while infrastructure choices moved slowly
Khan Academy formalized the distinction between one-way and two-way doors. Khanmigo launched under the experimental Khan Labs banner and remains a pilot with participating districts, making features reversible as evidence arrives; if tutoring proved terrible for a subject such as string theory, the team could turn it off after learning from actual use.
A contrasting one-way door was the backend migration forced by Python 2’s end of life on Google Cloud. Choosing Go and decomposing the monolith into services required careful analysis of developer productivity and server expense because reversing the language decision would be costly. Organization-wide decision records now document major strategic shifts; the framework came from Amazon, under “copy and customize.”
Moving early also meant building infrastructure that later became available elsewhere. Khan Academy began with a custom Google Sheet for labeling, created its own prompt-engineering interface, and is looking into regression tests for prompt changes. It is now rewriting a lot of code around LangChain for easier model swaps and JSON responses, while considering specialized external labeling products.
Jancipar’s defense of the debt was blunt: “Sometimes tech debt is good.” Hitting the GPT-4 launch mattered even if a cleaner architecture would have shortened aggregate development later, because early execution, Khan Academy’s brand, and its trust caused some observers to conclude, “Why would I even bother entering the space?”
7. Private reasoning pushes outward the model’s can–can’t boundary
Nathan tested that boundary with a boat-and-waterline problem involving subtle physical reasoning, then asked whether Khanmigo should disappear beyond subjects it could handle. Jancipar’s answer was to enable it broadly, disclose that it makes mistakes, and learn. Older university students should often recognize weak help; the governing question is instead, “Is this harming anyone?” and how can that harm be reduced?
For mathematical messages, Khanmigo can run a first call that asks whether the topic is math and, if so, privately works through the student’s answer, the exercise context, correctness, and likely misconception. That private tutor-like analysis is hidden from the learner and supplied to a second call instructed to tutor from it; Jancipar said this significantly improved performance.
Nathan connected this hidden self-critique to Tree of Thoughts, which explores multiple reasoning branches and selects among them, while noting that the token count can become multiplicative. Jancipar described chain-of-thought as an implemented technique but did not say Khanmigo was using Tree of Thoughts.
Context determines reliability. Exercise tutoring works best because Khanmigo receives the question, answer, and human-authored hints; the open-ended “Tutor Me: STEM” activity lacks a known answer and is weaker. One prospective fix is to solve the problem first through Wolfram or a Python backend, inject the result and steps, and only then begin tutoring.
8. Retrieval and cache design reduce both hallucinations and inference cost
Khanmigo’s embedded knowledge base began with questions the base model could not answer about Khanmigo itself. “AI Power,” for example, is a daily allowance capped at 200,000 tokens—high enough for ordinary use and intended primarily to prevent abuse, though cost is also a factor. When an incoming question clears a similarity threshold, the relevant product information is injected into the prompt.
An earlier retrieval experiment showed the failure mode: automatically suggested links could send an Algebra 2 learner to an MCAT course. The feature was temporarily removed because it was not valuable enough, although the team still wants reliable link insertion when a learner needs a particular Khan Academy resource.
A dedicated OpenAI instance supplies additional caching. OpenAI advised placing the largest shared prompt prefix first and moving user-specific variables toward the end; reorganizing prompts that way “significantly improved” the cache-hit ratio. Nathan proposed an NVIDIA GPT-4 Minecraft-agent analogy—save successful solutions as reusable modules—but Jancipar cautioned that the permutations of student questions first need to be understood.
9. Safety, memory, and transparency are inseparable product requirements
Jailbreak resistance comes from both GPT-4 and Khanmigo’s orchestration. User messages can be wrapped with reminders to remain on task, and every message passes through OpenAI’s moderation API for signals including sexual or otherwise inappropriate content. Nathan described the change from early versions to the current state as “night and day”; Jancipar said OpenAI had continued iterating, including on the June model, without claiming the broader safety problem was solved.
Personalization in production is still narrow: learners can choose simple, complex, or professorial reading styles, while exercise context may say whether they are unfamiliar, familiar, proficient, or mastered. It does not yet maintain the comprehensive learning biography implied by the long-term vision.
The roadmap includes opt-in interest extraction and continuity across learning threads. One student discussed sine, moved to tangent, and was frustrated when a new backend conversation again asked what she knew about sine: “I already told you about sine like 10 minutes ago.” To users, Jancipar noted, the journey should feel continuous even when the system starts a new thread.
Khan Academy also teaches users what the system is. Jancipar recalled that an AI course for teachers or students had launched, possibly in March, and support material feeds the knowledge base. The interface says, “Khanmigo makes mistakes sometimes—here’s why.” Students are explicitly told that teachers and parents can review their logs, a condition Jancipar sees as essential to family trust.
10. Voice and classroom orchestration broaden Khanmigo beyond chat
Text-to-speech ranks high because even simplified responses remain text-heavy. A seventh grader may read at a fifth- or second-grade level, turning a capable written tutor into an inaccessible one; Khan Academy is testing providers, multiple voices, and possibly unlockable voices as an engagement mechanic.
Visual input could let a learner scan homework in the mobile app or write mathematics directly on a touchscreen. Jancipar’s preferred intervention is not necessarily to interrupt the first mistake: after a repeated error, Khanmigo might ask why the learner chose that step, preserving productive struggle while drawing out the misconception.
Teachers already have a teacher-side system for lesson plans, differentiated learning plans, and insight into where students struggle; Jancipar said many teachers feel it saves meaningful time. Students have reacted positively to an always-available, nonjudgmental tutor, while learners whose first language is not English particularly value discussing difficult material with Khanmigo in their native language.
Marc Bhargava rejected a future of children learning alone at computers from grades one through twelve. The aim could be to reduce 10 raised hands to perhaps two and free teachers for relationships and project-based work. Future multi-user activities might group learners by shared misconception, let a class collectively build a story, or have AI facilitate and grade debates.
11. Efficacy and equitable distribution remain the decisive tests
Early evaluation asks whether learners spend more time on Khan Academy and cover more material; even story creation may strengthen reading comprehension. Jancipar nevertheless called standardized assessment the “Holy Grail,” and gave no timeline for a Khanmigo efficacy result.
Khan Academy’s precedent is its work with NWEA’s MAP Growth test. Scores were ingested to assign individualized goals across areas such as measurement and geometry, students practiced at their own pace, and later testing showed statistically significant gains versus learners not using Khan Academy. The proposed Khanmigo study would compare ordinary Khan Academy users with a cohort receiving Khanmigo alongside the same content.
Khanmigo is not yet free; Jancipar said it remains significantly cheaper than a real tutor. Districts pay per student per month, while Khan Academy prioritizes schools with more students receiving free or reduced-price lunch, discounts access, and recruits local corporate sponsors. A contemplated membership ladder would let higher-paying supporters offset licenses for historically under-resourced learners; Nathan’s individual route was a recurring $9 monthly donation.
Jancipar’s final metaphor is the “Swiss cheese gap”: old Khan Academy discovery was roughly O(log n) search through possible weaknesses, whereas an AI tutor might pinpoint the missing skill and return a learner to grade-level work. Videos gave everyone “Sal as the world’s teacher”; Khanmigo could add personal attention at scale. His conviction stayed qualified: “I think they’re going to show that this can really help students, and I think it can help change the world.”