Pioneers Insight Method Research Author
E197 | An Anime Producer’s Journey Through Japan, and the Multimodal Race Among Seven Models
Back to Episodes

E197 | An Anime Producer’s Journey Through Japan, and the Multimodal Race Among Seven Models

Summary

  • Global anime demand is growing by more than 10% a year, while the number of potential suppliers capable of matching top-tier anime aesthetics may be fewer than 10, making it easy for a new project to stretch from launch to visible results over five years. A 12–24-episode season takes about 3 years of production alone, and top studios typically have 2-year backlogs; an episode costs about $400K, with high-end productions exceeding $1M, yet remains an order of magnitude cheaper than Western animation. “The market interest is here”; what is truly scarce is talent, management systems, and aesthetic fit.
  • Video generation has entered a “stampede,” but the industrial hurdle is not whether a model can generate an image—it is whether it can reliably and controllably fix the final 5%. If 10 frames are each 95% correct, the estimated probability that all 10 are correct is roughly 0.95¹⁰; if animation supervisors still need to inspect every frame, moving AI from 90% to 95%, then to 100%, may not be faster than having a person draw it directly. “Everything comes down to what counts as ‘enough’(够).”
  • Traditional studios initially focused on in-between-frame generation, but it has yet to become truly usable; director assistants and ordinary static backgrounds have emerged as more practical entry points. Feeding a director’s past scripts, storyboards, and revision notes into ChatGPT could, even if it only handles 40%–50% of the initial screening, raise the number of critical points the director can process each day from 10 to 15; background generation works, but backgrounds are usually not the biggest bottleneck in animation capacity.
  • AI-native teams can bypass traditional pipelines and cut costs, but they also sacrifice the exaggeration and aesthetic design that make anime valuable. Kaka Creations, a team of about 10, used live-action motion capture and AI style conversion to make a roughly 30-minute animation it claims is 95% AI-generated; Tianyu rated it “7 points, with 6 as passing.” The efficiency is real, but the movement is often stiff: “If I genuinely liked this style, why wouldn’t I just watch a live-action movie?”
  • Chinese and US video models have yet to establish a clear generational gap; competition is centered on update speed, stability, generation speed, pricing, interaction, and vertical aesthetics. Tianyu believes Kling, Hunyuan, and Vidu are “not inferior to US models” in generation quality, with Kling particularly strong on anime keyframes; Pika, Runway, and Luma may be stronger on effects. Veo 3’s addition of sound, lip-sync, and audio-visual synchronization is striking, but he expects other vendors to follow quickly.
  • Current models have largely solved hollow eyes and extra or missing fingers, but remain constrained by duration, narrative continuity, and compute costs. Standard one-shot generation is about 10–20 seconds; 20–30 seconds is already considered long and relatively stable, while anything over 1 minute requires genuine story comprehension. Mid-tier consumer subscriptions often run out in a week, while an industrial pipeline could jump from 20–30 experiments a day to 500, making “tokens, cost, and efficiency” immediate hard constraints.
  • Voice generation is approaching human quality and music can already produce usable results, but the closer AI gets to replacing people, the greater the commercial-ethics and talent-supply risks. Many well-known Japanese voice actors have publicly opposed using their voices for training or allowing AI imitation; automating in-between frames could likewise remove the first rung on the ladder for newcomers to practice and advance. The ideal outcome is not a poorer creative ecosystem, but AI as a cheaper new pigment, enabling “a hundred schools of thought and a hundred flowers to bloom.”

Deep dive

1. Azuki’s move from NFT avatars into anime first runs into a five-year timeline to results

  • Tianyu explained that Azuki emerged from the 2022–2023 Web3 wave as a brand combining NFTs with anime; after moving from an engineering role at Google Brain into content, he now mainly oversees content development, especially anime production.

  • Azuki prefers to build IP through series, because comics and serialized shows remain the most direct entry points for anime fans discovering new IP. But in execution, a short film or feature may actually move faster than a season, creating the industry paradox that “a movie is slightly easier to make than a series.”

  • A season typically runs 12–24 episodes and takes about 3 years of production alone. Top studios commonly have production lines booked more than 2 years out, so a new project starting from zero “very easily becomes a 4- or 5-year timeline,” with results potentially not appearing until year 5.

2. Demand growing 10% a year meets a pen-and-paper workshop supply chain

  • The generational shift on the demand side is clear: people born in the 1980s and 1990s grew up with anime and have entered the consumption mainstream, while Gen Z and Gen Alpha are taking over behind them. Including streaming, video platforms, mobile games, games, and collectibles, Tianyu estimates that the global anime market is still growing by more than 10% a year.

  • Supply remains highly labor-intensive. An episode’s credits often list a long chain of outsourced specialists, and each company may in turn employ hundreds of people. Even so, roughly 30%–40% of production is still done on paper, while the computer-drawn portion is also “people drawing by hand.”

  • Tianyu’s analogy is that Japan’s animation industry looks less like a highly integrated factory than a small ramen shop shaped by “craftsmanship and artisan culture.” Digitalization, management capability, and talent scale limit expansion—not a lack of orders.

  • An episode costs about $400K, and can be cheaper when an existing manga and streamlined pipeline are available. High-end productions can exceed $1M; an episode of Arcane, with a budget in the $10M range, would be extremely unusual in Japan, while Western content often costs more than an order of magnitude more.

3. Global capacity exists, but fewer than 10 suppliers truly match top-tier IP aesthetics

  • Japan’s lead comes not only from having more production workers, but also from long-term government support, a mature system shaped by cross-pollination across the industry, and a deep talent base built over time. China has seen sustained investment from Bilibili, Tencent, and others, with quality studios “springing up everywhere,” but the industry as a whole is still catching up in maturity.

  • Tianyu cited Chinese animation titles including The Memory Bureau, Mutant Hero, Super Cube, and Ne Zha 2 earlier this year. He believes Ne Zha 2 may be a breakout moment for China and expects China to begin closing the gap with Japan in both production quality and efficiency.

  • The number of potential partners capable of taking on top-tier anime projects can be “counted on your fingers—probably fewer than 10.” China, South Korea, and Southeast Asia are not doing poor work, and Europe is strong as well, but Europe leans toward artistic expression, while the US has spent decades developing visual languages associated with Marvel, DC, Disney, and Pixar.

  • Tianyu believes “more than 50% of content production may depend on a team’s taste in a work.” Azuki was born in the US, but its visual DNA is close to Japanese anime. Its art director, Steam Boy, was previously Blizzard’s director of character design and designed the first Overwatch characters, making it surprisingly difficult to find production capacity with the right sensibility in the US itself.

4. Anime shorts reduce runtime, but not costs proportionally

  • Animation is not captured continuously by a camera; movement is drawn one image at a time. A few seconds of action may require 15–30 drawings, and a full episode can accumulate thousands of hand-drawn keyframes. A live-action short can generate a large amount of footage quickly; animation cannot eliminate the production of each frame.

  • At roughly 10 frames per second, 30 seconds of motion could require 300 drawings. But cost cannot be inferred from frame count alone. A person sitting and drinking coffee might be extended to a minute with 10–20 drawings; change the action to running, walking a dog, or dancing, and the difficulty of both the movement and each drawing rises sharply.

  • As a result, the per-second cost of an anime short can be higher than that of a live-action short. The real budget drivers are movement intensity, performance detail, effects, and the degree of aesthetic exaggeration the director wants to preserve—not runtime itself.

5. Traditional studios are all watching AI, but still lack a stable in-between-frame solution

  • Japanese mainstream studios rarely promote AI publicly because the subject is highly sensitive among artists. But after speaking with studios on this trip, Tianyu found that “basically every anime studio is looking at AI,” seeking productivity gains in scripts, character design, storyboards, first and second key animation, in-between frames, backgrounds, music, voice acting, and post-production.

  • The most obvious entry point is in-between frames. An animator first draws 3 key states—the hand touching a coffee cup, the cup reaching the mouth, and the moment after the sip—then fills in the movement between them. Keyframes are considered creative labor; in-betweening is relatively repetitive and is often the first job through which newcomers enter the industry.

  • Technically, the task is to give a model the keyframes at both ends and have it generate the intervening video. New breakthroughs appear in papers almost every 1–2 months, and R&D teams at Bilibili, independent US research groups, Chinese universities, and domestic companies have all shown promising results. But the studios Tianyu visited came away from testing with the same conclusion: “not stable enough.”

  • That “enough” requires the model to be both credible and aesthetically pleasing. If a jacket fold, glove texture, pattern on a cup, or patch of light flickers in and out, viewers may not only see a continuity error; they may interpret the mistake as a plot device, because every mark in anime is generally assumed to be intentional.

6. The final 5% of errors compounds and consumes the first 95% of efficiency

  • Tianyu’s production arithmetic is straightforward: suppose AI generates 10 frames, each reaching 95% accuracy, but with errors appearing in different places. Multiplying the probabilities frame by frame gives an estimated probability of roughly 0.95¹⁰ that the entire sequence is fully correct. If animation supervisors and key animators still have to inspect every frame, the time saved can be consumed by rework.

  • When 10 interns make mistakes, the supervisor can explain the problem collectively, and the next batch of drawings will generally improve. A generative model, even with masking, may damage other areas when asked to remove a single fold. “The process from 90% to 95%, and then from 95% to 100%, is genuinely not necessarily faster than a person.”

  • Hollywood explosion scenes illustrate what controllability means in practice. Directors may specify the size of the explosion, the color of the smoke, and the direction in which debris flies. Tianyu noted that James Cameron may simulate each explosion hundreds of times. For AI to enter film and television at scale, the industry will need sustained investment in this kind of fine-grained creative control.

  • The standard of “enough” also changes with the context. If someone uses a deceased relative’s photo or voice to generate a 10-second memorial clip, “something is better than nothing,” and imperfections may not matter. Once the system enters industrial content, continuity, editability, and aesthetic intent become hard requirements.

7. AI-native animation can pass, but motion capture struggles to replace anime exaggeration

  • A second class of startups is abandoning the old pipeline altogether and adopting the strategy: “We do whatever AI is capable of.” Kaka Creations, a team of about 10, produced a roughly 30-minute short it claims is 95% AI-generated. Tianyu gave it a score of 7, with 6 as the passing grade.

  • Rather than asking AI to draw every movement from scratch, the team filmed live-action performances such as picking up a coffee cup and taking a sip, then used AI to convert them into an anime style. The cost and efficiency advantages are real, but the expressiveness is “frankly still far behind,” and viewers can easily tell the work was made by AI.

  • The problem with motion capture is not that it is unrealistic, but that it is too close to reality. Real human smiles, body movements, and fight choreography have physical limits; anime stretches the smile wider and turns the eyes into lines. Using Doraemon as an example, Tianyu said the motion-capture conversion was “not exaggerated enough, not artistic enough, and not fun enough.” If the audience only wants realistic movement, why not watch a live-action movie?

8. Background generation has a clear use case, while a director assistant may be closer to the high-value opportunity

  • Hongjun proposed first taking a photograph and then having AI convert its style. Tianyu believes static backgrounds are the easiest application to make work and are less likely to break continuity. He cited Netflix’s PLUTO, which has publicly said it introduced AI generation into background production. But backgrounds are generally produced in parallel with character movement and are not animation’s primary bottleneck; environments at the standard of Makoto Shinkai are also far beyond ordinary generation.

  • After seeing limited results from keyframes and backgrounds, one traditional studio fed a director’s past storyboards, suggestions, and scripts into ChatGPT and asked it to simulate the director’s evaluation of a new project. The internal tool did not draw animation directly, but it won the director’s approval.

  • An animation director has to evaluate storyboards, scripts, color, movement timing, and plot, and cannot personally correct every frame. So-called “off-model animation” often results from a downstream team dropping the ball, not from the director lacking the ability to fix it. The practical constraint is the enormous workload and limited attention available.

  • Tianyu believes that even if AI can help a director process only 40%–50% of the critical points, it would still be useful. If the number of points a director can revise with full concentration each day rises from 10 to 15, that would be a huge success. Hongjun summarized the effect as “a 30% improvement in quality” and believes this path has better prospects.

9. AI’s greater value may be creating visual languages that were previously too expensive to draw

  • The Japan trip made Tianyu more cautious about AI deployment in existing pipelines, but more excited about small AI-native teams. The question is not only which labor can be automated; it is also that “things that were completely impossible to make before can now be made.”

  • Complex clothing is one example. 2D hand-drawn animation rarely puts characters in elaborate patterns, ornaments, and bells, and it is difficult to draw medieval armor frame by frame. That does not mean the designs are unattractive; animators simply cannot afford the cost of keeping them moving.

  • Coloring appears to be an obvious automation target, but it reveals the subtle relationship between technology and aesthetics. Much anime coloring resembles the paint bucket in Windows Paint, filling enclosed areas with color. It is repetitive labor, but also a technical condition that helped shape the visual language of existing animation.

  • Tianyu compared this with Greek sculpture. Because pigments were difficult to preserve over time, later neoclassicism came to treat whiteness as an aesthetic feature. Once chemicals and plastics matured, they enabled new production systems behind Transformers, Doraemon, and collectible figures. The most exciting potential of AI is not copying old creative ideas, but opening new forms without erasing the contributions of individual artists.

10. Azuki uses AI to make avatars more engaging, while deliberately preserving handmade scarcity in the collectibles

  • Tianyu’s day-to-day work includes writing stories, designing characters, following production, fundraising, promotion, and resource integration. Azuki did not begin as a comic or novel, but as a set of NFT avatars. The team therefore continues to experiment with how to make existing avatars “come alive” and add interactivity without breaking the logic of collecting.

  • But Azuki’s original avatars were not generated with AI. The collectible value of an NFT comes not only from image quality, but also from handmade production, supply control, and scarcity. Tianyu emphasized that AI can solve production, but does not automatically solve promotion, commercial value, or whether audiences will be moved; those still depend on a director’s and creator’s understanding of cultural works.

11. Video-model competition has shifted from isolated breakthroughs to continuous catch-up

  • Tianyu dates the start of the “stampede” to roughly 7–8 months ago, but does not see Sora as the only marker. Compared with the high expectations and controversy surrounding Sora’s release, progress from Kling, Pika, and Runway at several major inflection points may better represent the industry’s shift into continuous catch-up.

  • Competition is concentrated in release frequency, stability, speed, and prompt comprehension. Google Gemini first demonstrated text-based editing of a single image; weeks later, ChatGPT launched a similar capability and triggered a surge of Ghibli-style generations. By comparison, Midjourney and Stable Diffusion historically had “not particularly strong” command of textual logic.

  • Luma showed late last year that it could take a defined starting point and endpoint and automatically fill in the movement between them. Tianyu believes China’s Kling may have appeared almost simultaneously, with higher quality in anime-style keyframe interpolation. Chinese models including Hunyuan, Kling, and Vidu are “not inferior to US models” in generation quality, while terminal experience, speed, and pricing may be better.

  • Markets also show different aesthetic biases. Chinese teams are more familiar with anime and may naturally favor it in their training data and product decisions; US teams such as Pika, Runway, and Luma may be stronger on effects. Tianyu currently sees no “true generational gap” between any 2 leading companies.

12. Veo 3 adds sound, but long-form narratives still jump like dreams

  • Google released Veo 3 during I/O, adding sound, lip-sync, and audio-visual synchronization to video generation. Tianyu acknowledged the achievement but believes the model may not have created a deep moat; similar capabilities could appear in other products quickly.

  • Mainstream video lengths remain in the 10-, 15-, and 20-second range, with 20–30 seconds already considered long and relatively stable. Beyond 1 minute, a model must do more than extend an action: it must understand context and the story line. “No one wants to watch a person pick up a coffee and drink it for a minute.”

  • Hongjun once asked Veo 3 to generate a squirrel and a cat running across a hillside, through woods, and over a bridge before arriving at a mountaintop with a rainbow and wind. The model included every element, but substituted scene cuts for continuous running, while objects underwent astonishing deformations. The two described the result the same way: “It felt like a dream.”

  • Local quality has improved dramatically. Characters’ eyes have largely escaped the vacant look of early models, and an extra or missing finger has become an occasional minor bug rather than something that must be checked every time. But subscription quotas remain tight: 2 or 3 ideas tested 5–10 times on each platform already means 20–30 attempts a day, and mid-tier plans often run out in the first week. Industrial production could require 500 attempts a day.

13. Voice and music have crossed the technical threshold, but now collide with rights and livelihoods

  • Tianyu believes the generation quality of most cutting-edge voice models is already “indistinguishable from a real person.” Japan’s voice actors have their own association, and in recent months many well-known actors have publicly opposed AI, refusing to let their voices be used for voice training or imitated by AI, because their voices, voice training, and performances are their livelihoods.

  • It is difficult to keep describing AI merely as a “tool.” If a system allows Hongjun to write scripts in the future without recording his own podcasts, the subjective shock is already close to replacement. A voice actor contributes not only physical sound, but also character performance, audience reach, commercial pull, and creative input.

  • Tianyu believes music can also be generated. As for expressiveness, he thinks half the answer has to come from listeners themselves. The combinations of major and minor keys and rhythms that humans find pleasing are not infinite; predecessors and music theory have already cataloged them extensively, so it is not especially difficult for a model to understand “things humans find pleasant.”

  • Hongjun once felt that Suno’s pop music sounded too much like disposable bubblegum. People working in the field explained that the platform avoids training on or copying contemporary top songs. Classical music performs better because many works are already out of copyright and the data is more open. If technology can truly copy Jay Chou’s voice to write songs, whether new artists can still earn returns—and whether humans will still create good new music in the future—becomes a structural question.

14. If automation removes the entry-level rung, long-term capacity may actually decline

  • Automating in-between frames appears to be the least controversial path: people still create the keyframes, while AI takes over repetitive labor. But in-betweening is also where newcomers accumulate practice and gradually become key animators and masters. If that rung disappears, the industry will lose the next generation of talent before it has entered the field or been recognized.

  • Tianyu’s analogy is that “if only the top half of a ladder has rungs you can grab, there’s no way to climb it.” With the industry already facing a large labor shortage, a short-term productivity gain that destroys the training mechanism could mean that animation “may actually go backward over the long term.”

  • The ideal path looks more like the chemical industry lowering the price of pigments. When tools become cheaper, painters do not disappear; more young people are drawn into the field. About 15%–20% of Tianyu’s own work now involves technical judgments about time, difficulty, reproducibility, and scaling costs. The team also tried training a model, but the results were poor, so it decided to “let the horse run ahead.”

  • Both believe hybrid talent will become more important. Technical teams need to understand creative standards and the final 5% of errors, while creators need to understand model limitations. Tianyu says his “left brain fights with his right brain every day,” while Hongjun admits he is more pessimistic about the future of AI and humans. Their final point of agreement is that this generation’s choices will shape the long-term relationship between technology and art: “We are the ones who have to write the answer.”