Pioneers Insight Method Research Author
142: Zebra CPO Xiu Jiaming: How an AI Product That Teaches Proactively Came Into Being
Back to Episodes

142: Zebra CPO Xiu Jiaming: How an AI Product That Teaches Proactively Came Into Being

Summary

  • Zebra Spoken English is not an old course with a chatbot bolted on; it is a proactive-teaching product driven by AI across the frontend and backend, UI, models and courseware. The product uses a three-way split screen with an AI English teacher, the child’s video feed and interactive courseware; each class runs 25 minutes. Each unit has 4 classes, each level has 96 classes, and there are 6 levels in total. Xiu Jiaming calls it the group’s first “full-stack AI product”; the AI Agent is only one component.
  • The truly expensive part is not calling the model, but compressing its freedom into the boundaries of deliverable teaching. The project began in August 2023, with roughly RMB200M ($28M) invested over more than two years and a team of about 200 people; at current capacity, a new level takes about a year. The team has to “preserve its flexibility while keeping it within defined boundaries,” a process Xiu calls “wrestling with the AI model.”
  • Children’s spoken English is a market where supply has long been scarce, while demand has been understated by the existing product format. Xiu observes that before roughly KET level—or before roughly PET level in first-tier cities—children can learn through various methods. After that, many hit a ceiling and are still cramming spoken English when they graduate from university to study abroad. His view is that “every child who learns English has this need.” Cars were once treated as status consumption for a small group as well: “Without the product, people don’t have the need.”
  • Zebra differs from ordinary AI chat companions because it has a proactive, year-long teaching plan and child-level long-term memory. Every child has an independent vector database recording their language, behavior, level, habits, interests and even their pet’s name. The company’s proprietary model handles classroom teaching, while DeepSeek is sometimes used to organize post-class information. Each class requires more than 100 valid speaking turns; the goal is not to maintain the current level but to raise ability level by level according to plan.
  • The strongest evidence so far is still process data and long-term beta testing, not publicly demonstrated end-state learning outcomes. The product has accumulated more than 7000 test classes; the earliest testers have been learning for more than six months, with the longest-running user completing 26 weeks. The team is launching a “spoken-English strength” metric built from accuracy, fluency and richness, and believes 4 completed classes are enough to estimate a child’s level with reasonable accuracy. Xiu says “every class is an exam”; at L4, the team will use a KET speaking mock exam for external validation.
  • Commercially, the team prices it as a learning product rather than a cheap AI tool: RMB3600 ($500) per level per year. During beta, it sold for roughly RMB1500–2000, and parents cared more about “what result will be delivered in how many days” than whether the teacher was human or AI. The team believes a user base below 100,000 would not create much capacity pressure. But the launch includes only 1 level, with the second to follow soon and roughly 1 new level added each year, so content production will directly constrain expansion.
  • Model upgrades are more likely to create a supply-side tailwind for Zebra, but the moat depends on whether teaching experience can keep translating into outcomes. Xiu welcomes stronger models and real-time AIGC courseware, and says the current fixed courseware could eventually be removed once the technology matures. But he stresses that in children’s education, “data is less persuasive than experience”: a model cannot determine from general-purpose training data alone what to teach a particular child first or how far to simplify. The team’s product standard is: “If a teacher or a person can do it, we don’t do it.”

Deep dive

1. Zebra rebuilt the 1:1 foreign-teacher class as an AI product from the ground up

  • The product uses a three-way split screen: a 2D AI English teacher, the child’s own live video feed, and interactive animated courseware that carries the main teaching load; each class lasts 25 minutes.
  • The curriculum has 6 levels, with 96 classes per level and 4 classes per unit. It looks like a human 1:1 class on the surface, but AI is integrated into and drives everything from the frontend and backend to the UI.
  • Xiu defines it as a “spearhead” product aimed at solving a narrow problem: start with children’s spoken English rather than building outward from a general-purpose education assistant.

2. Children’s speech data collected since 2014 became a foundational asset for the LLM product

  • Yuanli Technology set up its AI Lab in 2014, initially focusing on adaptive algorithms, computer vision, ASR, TTS and assessment—not large language models.
  • Zebra’s existing products accumulated a large volume of children’s speech, which helped train recognition and English-assessment models better suited to children’s age groups. Xiu calls this “previous-generation” technology, but says it remains the foundation of classroom usability.
  • Zebra Spoken English was launched in August 2023. Before that, adding large-model dialogue to the old products was merely “a product idea from the previous era, with some AI features added.”

3. The team did not eliminate hallucinations; it confined creativity within teaching boundaries

  • Xiu believes model capabilities were already broadly sufficient when the project began. Subsequent model advances mainly improve presentation and efficiency; the core challenge from day one was determinism.
  • The team spent more than two years “wrestling with AI models.” His counterintuitive view is that “hallucination isn’t necessarily a negative term”; sometimes it is precisely the model’s creativity.
  • When a child mentioned a gift, the AI companion volunteered, “I’m jealous.” The emotional response delighted the child, while the model filtered vocabulary according to the child’s level to keep the exchange broadly understandable.

4. A 25-minute class is broken into more than a dozen small goals so the Agent stays in charge

  • Each class is divided into more than a dozen segments of 2–3 minutes, each with a defined objective and achievement data. The AI knows in advance what must be completed, keeping the child’s language directed toward the teaching outcome.
  • This is not a companion bot that lets children chat freely. It is a proactive Agent that advances tasks alongside the courseware and pulls the child back on track when necessary. Every 25-minute class is validated as an independent product.

5. Safety comes from layered constraints across training, testing, monitoring and human review

  • Training includes data cleaning, classification, reinforcement training, adversarial “bug hunting” and rewards for positive examples, teaching the model which expressions it must not use.
  • The large test set used before launch is continuously updated with real-world data. After launch, Yuanli’s large model supports real-time monitoring of erroneous content and sensitive words, while humans continue to review and label cases.
  • Real-world problems flow back into the models and policies. The team is not trying to prove model safety once and for all; it continuously narrows the error space one class at a time.

6. The scarcity in spoken English is not knowledge, but a steady supply of suitable real-time conversation

  • Listening, reading, vocabulary and grammar can be self-taught or taught in batches by a single teacher. Speaking requires a real-time partner who understands the learner’s level and provides abundant opportunities to talk.
  • A non-specialist native speaker can chat but cannot design a teaching path or systematically break through individual knowledge points. The supply of Chinese teachers who are both professionally trained and sufficiently fluent is also limited. As a result, many children stall around KET or PET.
  • Cheng Manqi asked whether writing posed a similar output problem. Xiu’s distinction was that writing allows learners to stop and look up words or imitate model essays, so a 3-month cram can produce visible gains; in live speaking, “once you stop, the conversation can’t continue.”

7. “Every child who learns English” is a demand judgment, not a verified market figure

  • Xiu does not want to put a definitive number on the market size, but believes “every child who learns English has this need,” because the need for the “speaking” component of listening, speaking, reading and writing requires no separate proof.
  • His analogy is the automobile: early cars were slow and looked like a status purchase for a small, wealthy group, so the demand appeared limited. Once the product spread, its broad usefulness became visible: “Without the product, people don’t have the need.”

8. Chat companions can maintain a level; proactive teaching must decide what to learn next

  • The ceiling for passive AI chat is often maintaining the learner’s existing level. Improving ability also requires deciding the increment, practice volume and choice of materials—questions that cannot be left to a general model’s random improvisation.
  • Each level in the Cambridge system is tied to vocabulary and to the materials, daily life and cognitive range a child can encounter. An ordinary model understands neither this progression nor the long-term, year-long teaching plan.

9. Long-term memory lives in each child’s vector database, not in the model parameters

  • Every child has an independent database storing what they said and did in each class, along with long-term information about their level, habits, state, interests, family and even their pet’s name.
  • A new expression triggers relevant historical records, which are combined with current information to generate feedback. Classroom dialogue is handled mainly by the proprietary model; DeepSeek sometimes handles post-class reports, information organization and dimension-level summaries.

10. Teaching ability is woven into scaffolding that steps difficulty down level by level

  • When a child cannot answer a wh-question, the AI adjusts to the child’s level by switching to a choice question or yes/no question, providing a hint, leading the child through the answer, and then returning to the original goal. The key judgment is whether to step down 1 level or 2.
  • Analogies must also draw on prior knowledge. If a child has learned that come becomes came, the AI can use that pattern to guide them toward inferring that become becomes became rather than simply giving the answer.
  • Difficulty boundaries cannot be hard-coded. If a child spontaneously brings up the Golden State, Ultraman, Elsa or Eggy Party, the AI must engage with the general knowledge and then steer the conversation back to the lesson. The rules adjust dynamically based on whether the child initiated the topic and understood the exchange.

11. A warm personality serves expression, rather than merely delivering emotional value

  • The proprietary model incorporates pedagogy, psychology and linguistics, and is required to respond positively at all times. It cannot simply say “You’re great”; it must identify which sentence or part of the child’s recent response was good.
  • Xiu stresses that “good” is not the end goal. The aim is to lower the emotional filtering threshold around speaking a foreign language, keeping children comfortable and pressure-free so they are willing to continue producing language.

12. Children’s speech infrastructure determines whether the class is actually usable

  • The team simulates background noise, distance and speaking direction in home environments, handling situations in which children do not face the microphone or the surroundings are noisy.
  • The system can filter out adult voices. In households with 2 children, it also remembers the face and voice of the child taking the class, minimizing interference from younger siblings.

13. More than 100 valid speaking turns per class explain why the progress is slow and intensive

  • In 25 minutes, the child must produce more than 100 complete utterances, or roughly 4 per minute. Xiu says this puts the class into a state of “very intense English learning,” with the child continuously completing tasks alongside the courseware.
  • Progress is usually not linear. The first month can be relatively fast as children become familiar with the class and gain the confidence to speak; after roughly 3 months, they may hit a plateau before breaking through over 2 weeks to 1 month.
  • That is why each level is designed to last a year. Progress in speaking is “visible in small details,” unlike writing templates, which lend themselves to short-term drilling.

14. Animation and Easter eggs are temporarily driven by strategy combinations; real-time AIGC remains constrained by latency

  • In the “draw Mom” segment, the system assembles a scene based on the child’s description. Children are often shocked by the result but answer “no” verbally. This is not real-time model generation yet, but highly flexible composable courseware.
  • Easter eggs must trigger in the right context. After a child said they did not want a birthday party and became dejected, the AI replied that “even without a party, you still get a present,” then brought up a gift animation. Not every child will encounter it.
  • Xiu hopes to generate courseware in real time based on each child’s interests, but “talking about Ultraman and then waiting 5 minutes for it to be drawn” does not work. A children’s class also cannot rely on 2 faces talking to each other; asset-generation speed is a hard constraint.

15. The AI English teacher has developed an independent persona, but the team never designed it to “be like a human”

  • Cheng Manqi worried that children might confuse a personable AI with a real human. Xiu replied that the product uses a 2D character and was never designed around “how many percent human” it should look.
  • His analogy is: “A dog is a dog.” A pet can have human qualities and form an emotional bond without being mistaken for a person. The AI English teacher likewise occupies a new position in the child’s cognitive system.
  • Children form a mental schema through its consistent responses, learning roughly what they will get back when they say something. Xiu believes children know from the beginning that it is not human and do not mistake it for one; they more readily accept a new form of existence with a degree of human-like character.

16. Defiant children must be brought back to the lesson within 2 turns, while novelty fades with long-term use

  • When a child goes off topic, the AI must return to the courseware within 2 turns. If the topic is too far afield or involves sensitive content, it does not answer; when necessary, it firmly reminds the child, “We’re in class,” and uses a simple task to pull their attention back.
  • These situations were rare in testing. By the fourth or fifth class, battling the model was no longer novel; many children had already used DeepSeek or Doubao, so their attention in class shifted toward the English task.

17. More than 7000 beta classes exposed small gaps in contextual responses, not catastrophic errors

  • The beta has accumulated more than 7000 classes. The earliest testers have been learning for more than six months, and the longest-running user has completed 26 weeks. The product definition has not been fundamentally overturned, but every class is reviewed and produces local iterations.
  • On Teachers’ Day, a child showed Jessica a card and flowers. The AI smiled and thanked the child; the response was semantically correct but “not excited enough.” The team then added natural-calendar strategies and Easter eggs for Teachers’ Day, Christmas, Chinese New Year and other occasions.
  • Early on, one child had to read the same sentence 5 times and was close to breaking down, while the AI maintained “cold patience.” The team responded by adding round limits, segmentation, difficulty reduction, hints and material replacement.

18. Measurement is shifting from number of speaking turns to “spoken-English strength”

  • The internal “spoken-English strength” metric has 3 major dimensions—accuracy, fluency and richness—with each broken into 2 sub-dimensions and aggregated through objective data and formulas.
  • More than 100 utterances in a single class can provide a rough level estimate; after 4 completed classes, the estimate becomes more accurate. Xiu says speaking “doesn’t really need an extra exam,” because children have no crutches: “Every class is an exam.”
  • If the team lacks sufficient confidence in its own system, L4 is planned to deliver KET-level ability and include a KET speaking mock exam, using an external test to validate the outcome.

19. Price, capacity and teaching experience together define the business’s expansion limits

  • The course does not allow repeat learning. Each unit contains 2 dialogue classes, 1 expression class and 1 integrated-application class, plus pre- and post-review; knowledge is also revisited cyclically. The team believes 2 classes per week is already close to the upper limit for most families, and repeating a class would only waste 25 minutes.
  • Payment is calculated by level, at RMB3600 ($500) per year. Beta pricing of roughly RMB1500–2000 was well received. Parents care most about the delivered result and will directly ask for the child’s feedback after a trial.
  • The roughly 200-person organization spans curriculum development, animation and interaction, product, engineering, audio/video infrastructure, the AI Lab, labeling and quality control, with cumulative R&D investment of about RMB200M ($28M). Xiu says reliability falling from 99.9% to 97% would be unacceptable because “3 out of 100 people would have encountered an educational error.”
  • The launch includes only 1 level, the second will follow soon, and roughly 1 new level will come online each year, expanding from middle age groups toward both ends. Against general-purpose model competition, Xiu’s moat test is that “data is less persuasive than experience”: knowing what to teach a specific child first and how far to simplify cannot be replicated merely because all the materials are public.