Pioneers Insight Method Research Author
AIGCode Founder on End-to-End Agents and AI Coding's New Battleground
Back to Episodes

AIGCode Founder on End-to-End Agents and AI Coding's New Battleground

Summary

  • AIGC Code founder 宿文’s view is that Copilot-style code completion is more likely to be Big Tech’s game, while startups’ opening lies in an end-to-end Autopilot spanning front and back ends. His reasoning is that completion is “the shortest and fastest translation of a foundation model, almost an afterthought,” while Cursor and Windsurf depend heavily on Microsoft’s VS Code IDE ecosystem, leaving “the most critical pieces” outside their control. He confirmed that AutoCoder is positioned as the world’s first front-to-back coding product and could be available for public use by the end of June if all goes well.
  • His contrarian read on model boundaries is worth noting: models are naturally good at “translation and transposition,” while long chains of reasoning “still do not generalize today.” Entrepreneurs who radically reset their plans every time V3 or R1 launches may, in his view, be signaling that “they do not know where the technology boundary is.” If he has not understood that “the foundation underneath cannot support a 100-story building,” he is wasting equity capital.
  • His verdict on Cursor is that it is “cashing in before Big Tech and AGI giants roll over the market,” reaching $70M ARR in 9 months by getting just slightly ahead of the foundation-model dividend. The same play is harder in China: “A competitor gives you code, you build a solution first … and soon the solution is taken away without paying a cent.” He sees Silicon Valley’s ecosystem as relatively healthier.
  • To him, vibe coding is “the difference between a gashapon machine and real retail”: you accept whatever the model gives you, with no control. It is only “an intermediate state in the exploration of the technology,” somewhat like treating the automobile as a toy when it first appeared. When the host noted that quality tends to deteriorate after 1.5 to 2 hours of continuous coding in tests, 宿文 attributed it to hallucinations and context-window capacity. He agreed with the host’s warning about “web revenue” or “vibe growth”: much of the token usage “has not reached real consumer users.”
  • The key demand-side data point: Lovable drew more than 27M visits in April based on a crawl conducted on April 30, putting it within one order of magnitude of GitHub, which has been maintained for nearly 20 years. Lovable’s 85% first-month paid retention is a CEO-reported figure, while its 3-month ARR reached $17M—better than 宿文 expected. But it is “not a competitor”: it only handles the front end, while 80%-90% of software engineering work revolves around databases and the back end.
  • The case for training a proprietary foundation model is that many product optimizations in the LLM era sit at the model layer, while fine-tuning alone leaves “many flywheels” unable to turn; companies must be able to touch pre-training. His response to the idea that the foundation-model table is already set: “Before DeepSeek, a lot of people thought they were already at the table … nobody stays at the table forever.” The path is to iterate toward the next architectural step along MoE, rather than wage a compute arms race on the same platform; owning a model also makes a cost-based price war an option.
  • Writing good code is a small AGI loop, and within its target use cases coding will become infra in 3 to 5 years: “As long as you dial up compute, the code can be completed.” Coding startups without their own models may eventually have to train one or be acquired. OpenAI’s $3B acquisition of Windsurf in May struck him as “a very natural outcome.”

Deep dive

1. From a Small-Town Exam Ace to COO of an Industrial Software Company: 宿文’s Two Earlier Chapters

  • 宿文 describes his path as that of a “typical small-town exam ace” from Huining County, Gansu. He entered Tsinghua University in 2008 and “was crushed all the way” once he got there. The host said he graduated first in his PhD class; 宿文 only said he worked hard for a period. After leaving school in 2017, he started in VC, then joined RootCloud full-time as COO in March 2021. The company sells B2B industrial software with million-yuan-level contracts, and its team grew from fewer than 10 people to nearly 200.
  • His biggest lesson from that startup was the commercial loop itself: software businesses inevitably involve solutions, direct sales, round after round of negotiation, and even tenders. His wording was deliberately restrained: “I’m not saying the business model is problematic. I’m saying it may not be what I’m good at or what I want to do.”

2. Bezos’s Epiphany: “Retail Is the Sexiest Business Model in the World”

  • At the AWS AI startup accelerator last October, he saw Bezos’s line that “retail is the sexiest business model in the world.” It hit him as an epiphany: how can everything he is doing today eventually generate the kind of rapid, retail-like scaling leverage?
  • His framework is supply-demand matching: know the boundaries of what you supply and of the demand you are serving; if the two can be connected well enough, the business scales much better. That yardstick ran through the rest of the conversation, including his criticism of vibe coding.

3. Why Coding: The Most Important First-Tier Corpus and the Agent Vanguard

  • The starting point was a structural shortage of code supply: user requirements change and expand, while code supply invariably lags business growth. LLMs turn tokens into code, compressing iterations from weeks or months to minutes or seconds.
  • When foundation models are not good enough, moving prematurely into other industries prompts domain experts to say, “This is completely unusable,” making it difficult to create either data or commercial resonance. People who write code or algorithms, by contrast, can use their own expertise to compensate for immature tools and get the two gears turning together.
  • Code may be the most important corpus in the first tier; among synthetic data, “code and video are the most important.” He also believes that until coding matures, other agents may not run well. In 3 to 5 years, coding will look more like infra—possibly part of the operating system of the LLM era.

4. 2024: Copilot’s Year — Sogou Input Method for Programmers

  • His one-line summary of 2024 was the arrival of the Copilot product line: in existing development workflows, it provides programmers with an assistive co-pilot and completes part of the code. The host compared it to Sogou Input Method for programmers, and 宿文 agreed that Copilot is an efficiency tool rather than a replacement for the programmer.
  • The interesting mismatch is that 宿文 himself has never worked as a programmer: “I haven’t used it since I left campus.” He does not think that matters for product positioning. The key is to jump to the final step and ask: “Nobody pays for code. Why would I buy 100 lines of code? I’m paying for the software and application.”

5. The Day-One Conclusion: Copilot Looks More Like Big Tech’s Game

  • There were 2 reasons. First, Copilot sits close to the foundation model: it is “the shortest and fastest translation of a foundation model, something you do almost as an afterthought.” Second, it depends deeply on the IDE ecosystem—whether through plugins, cloud IDEs, or a modified version of open-source VS Code. Cursor and Windsurf both followed this path, leaving “the most critical pieces” outside their control.
  • The consequence is dependence in both product iteration and distribution. About a month earlier, Cursor had reportedly clashed with Microsoft over wanting to modify VS Code functionality that Microsoft would not open up. “China still does not have its own IDE”; Huawei and others are trying, but remain well behind the global leaders.
  • The thesis was later validated. In May, OpenAI acquired Windsurf for $3B; Anthropic has its own product, and Microsoft has GitHub Copilot. Similar tools can also be found inside Baidu, Alibaba, Huawei, and Tencent in China.

6. Devin and Cursor: Star Power and “Cashing In Before the Wheels Arrive”

  • In April 2024, Devin raised $175M at a $2B valuation without a product. 宿文 was not persuaded by the PR at the time: the team highlighted 10 Olympiad gold medals but “never said what it actually wanted to build or how it would build it.” The later $500 subscription and GPT-4-based multi-agent approach “were not quite right”; judged by the results, “it is still not ideal today.”
  • He has no shortage of praise for Cursor: “With that much revenue, if you still say it isn’t a good product, you’re just lying to yourself.” It reached $70M ARR in 9 months. The mechanism is to get slightly ahead of the foundation-model dividend—“not too early, or you won’t survive to see it”—then ride the momentum into code completion. He stresses that this was not necessarily a bet on Claude 3.5 specifically, but on what LLMs would eventually make possible. Some model would reach the required capability sooner or later.
  • Why is the same play harder in China? 宿文 thinks the Silicon Valley ecosystem may be healthier. In China, “a competitor gives you code, you build a solution first … and soon the solution is taken away without paying a cent.” Changing programmers’ IDE habits is also neither a startup’s job nor the kind of innovation driven by a startup’s strengths. He self-deprecatingly added: “I can’t say that if I decided to do it, I would definitely be able to pull it off. That would be too self-important and arrogant.”

7. 2 Foundation-Model Milestones—and Chinese Pioneers That Fell Before Dawn

  • 宿文 sees 2 important foundation-model markers in 2024: DeepSeek-V2’s contributions in engineering and reinforcement learning, with its architectural contribution “possibly a high-water mark”; and Claude 3.5. The latter’s technical dividend helped drive the emergence of Lovable, Bolt.new, v0, and Replit agent around October 2024.
  • Beyond advances in network architecture such as MoE, the larger contribution was that the code fed into the pre-training and post-training stages became increasingly refined. Cursor users “voted with their feet,” and later used Claude more heavily.
  • The domestic story had another side. 宿文 recalls that when 奇绩创坛 discussed agents in the first half of 2023, there were already 5, 6, 7, or 8 teams working on AI coding. Before the 2023 Chinese New Year, some shifted toward hardware or other directions. He suspects those teams did not survive to Claude’s arrival or to the point when China’s market and ecosystem reached PMF. Whether these model building blocks appeared did not greatly disrupt them: products aimed at replacing programmers end to end had major bottlenecks then, have them today, and will have them in the future. They need their own “brain” for continuous iteration.

8. The Contrarian View: Major Adjustments After Every Foundation-Model Launch Mean You Misread the Boundary

  • Asked about a possible Claude 4 launch at the time, 宿文 said he was not waiting for it. If every foundation-model breakthrough causes a major reshuffling of the application layer, “something is wrong.” After DeepSeek-V3 and R1 launched, people rushed to reposition and chase the hot topic; that may “mean they do not know where the technology boundary is and have no ability to anticipate it.”
  • The host called this a contrarian position. 宿文 explained it from a capital-allocation perspective: once a company has leveraged equity capital and committed people and compute, discovering later that “the foundation underneath cannot support a 100-story building,” and is only suitable for a villa or a bridge, means resources have been wasted. He also cited 王朔’s comment on 于丹: “From far away, the moon looks especially beautiful; up close, it is just one brick.”
  • He summarizes the model’s natural strength as “translation and transposition”: a flattened conversion from form A to form B. “Ask it to think through a long chain of logic, and it still does not generalize today. That is the ceiling of the Transformer architecture.” Expecting it to independently design the database, middleware, microservices, front end, UI, and APIs is “making the technology do something it cannot.” The way to use it is to narrow the boundary and amplify its strength in all kinds of business translation.

9. From Copilot to Autopilot: Vibe Coding Is a Gashapon Machine, Not Retail

  • 宿文 did not directly adopt Lovable’s CEO’s framing of AI replacing physical labor and then machines surpassing cognitive labor. He instead describes the current generation of products as “manufacturing content”: they have not yet reached matchmaking, much less the ability to affect the physical world. The self-driving analogy also has limits. L4 remains far away, but code bugs can be verified and fixed without the human consequences of an autonomous-driving failure.
  • Vibe coding—introduced by the host as a term proposed in February by OpenAI co-founder Andrej Karpathy—is a neutral concept to 宿文: “You accept whatever the LLM gives you, and it is completely uncontrollable.” That is “the difference between a gashapon machine and real retail.” It is only an intermediate state in technological exploration, “a bit like treating the automobile as a small toy when it first appeared.”
  • The host noted that in tests late last year, problems tended to emerge after roughly 1.5 to 2 hours of continuous coding, though performance has improved. 宿文 said the main issue is still the large context: as the session continues, information quality is affected by context capacity. The host also mentioned that some products use workflows and folders to help locate bugs.

10. OpenAI’s $3B Windsurf Acquisition: Coding Startups Without Models May Have to Train or Sell

  • 宿文 sees the acquisition as “a very natural outcome.” If an AI coding startup has no proprietary model, “it either builds its own model or may ultimately be acquired.” Windsurf’s engineering work, particularly around the IDE, can help OpenAI make the product larger and move faster.
  • Asked why Big Tech does not simply build AutoPilot itself, he said it still has to tackle the software-architecture layer. He initially described models and software architecture as roughly 50% each, then added that models may eventually become the larger share. Today, the model contribution is slightly smaller because model capabilities remain insufficient and need to be supplemented elsewhere.
  • He expects Big Tech to build similar products eventually: “Then let them.” The competitive edge will come from early-stage industry organization, people’s understanding, execution, and iteration speed. “Only what you iterate quickly enough counts.”
  • Copilot has accumulated programmer users, use cases, and user profiles, while products aimed at non-programmers require a different design. Even the core model capabilities that need iteration are different. On current evidence, Copilot will not naturally extend into Autopilot.

11. AutoCoder’s Innovation: The World’s First Front-to-Back Product, with a Proprietary Model and Generative Software Architecture

  • 宿文 confirmed that AutoCoder is positioned as the world’s first front-to-back AI coding product. In software engineering, “the front end may account for only a small part; 80% to 90% of the work revolves around the database and back end.” Overseas star products mostly stop at the front end, partly because they rely mainly on Claude 3.5 and 3.7, whose foundation-model capabilities remain limited. At a Sequoia Capital meeting in the US, the coding discussion was also “limited to the front-end development experience, with no mention of the full stack.”
  • Making it work depends on 2 things: the model, including a proprietary model and a clear definition of its boundary; and a generative software architecture. “Human logic and the logic an LLM uses to write software are different.” The system cannot write end-to-end software in the same way a human does.
  • The host’s plain-English version was that a developer needs 5 production inputs—A, B, C, D, and E—while a front-end-only product provides just 1. AutoCoder aims to supply all 5. 宿文 believes the vibe-coding moment is “more or less” here, and AutoCoder has already entered closed beta.

12. Lovable Is Not a Competitor, but It Validates Demand: 27M Monthly Visits, Within an Order of Magnitude of GitHub

  • Lovable and Bolt.new are “not competitors” because they only handle the front end and are not comparable products, but their user profile is exactly the one he wants to reach. Cursor is actually farther away: its users would come to his product with problems he cannot solve, and they would not expose their code. The front-end UI segment has a first-mover advantage, so he would rather let those products run than compete through homogenization.
  • Lovable’s performance “exceeded expectations.” Its 85% first-month paid retention is a CEO-reported figure, and its 3-month ARR reached $17M. Its strengths include operations, rapid iteration on small features, and an early integration with Supabase. Its engineering team or starting point was probably 2 months later than Bolt.new’s, yet it grew rapidly. 宿文 has not used it deeply because “it cannot solve my needs,” although he may receive free tokens from the company and test its capabilities almost every day.
  • The strongest evidence came from a crawl on April 30: even though Lovable only handles prototyping and website creation, its April traffic exceeded 27M visits—“compared with GitHub, which has been maintained for nearly 20 years, it is only an order of magnitude behind.” 宿文 believes supply must exist before demand can be released. The next stage will become an ecosystem play involving software go-to-market, as well as the impact on cloud services, payments, and general software distribution.

13. Why Train a Foundation Model: “Nobody Stays at the Table Forever”

  • The positive case is that many product optimizations in the LLM era sit at the model layer: “If this thing is not in your hands, you cannot optimize the content you create.” Even with data available today, fine-tuning alone leaves “many flywheels” unable to turn, so a company must be able to reach pre-training and complete the loop.
  • His rebuttal to the idea that the foundation-model table is already set: “Before DeepSeek came out, a lot of people thought they were already at the table. After it came out, many of them left the table. The same is true today—nobody stays at the table forever.” The path is not to compete for compute and data on the same platform, but to “move toward the second step”: iterate beyond MoE through architectures such as MMoE, CGC, and PLE to learn knowledge more efficiently, address some hallucinations, and train in product-specific templates.
  • For now, 宿文 is optimizing only in his own direction. Turning generalization into a token-selling business is “not something I particularly want to do today, and this may not be the right moment.” He believes models will become very strong but is not in a hurry. Asked why his team can do better than others, he cited experimental data from “alchemy” and the sustained dirty work: “Every day, move one brick and see what it looks like.”

14. Writing Good Code Is a Small AGI Loop, but Foundation Models and Applications Are “Asynchronous”

  • The logic is straightforward: code is the vessel for model output. The result of a general agent may ultimately be “mostly webpages,” but webpages are still missing many pieces; the missing piece in the loop is end-to-end code. The host described it as “the first egg laid on the road to AGI.” 宿文 agreed it may be an efficient path, but stressed that it is too early to reach a conclusion and that the priority should be a product that solves real problems.
  • “Asynchronous” means that foundation models have not supplied capabilities matched to applications: “They are laying the foundation as if they were building a highway, while you are building a skyscraper.” Relying entirely on foundation-model upgrades also creates performance swings. From Claude 3.5 to 3.7, some products continued to rely mainly on 3.5—not only because of cost, but because part of the tuned output quality was lost on 3.7.
  • He acknowledges a coordination effect between Big Tech’s foundation models and applications, but does not think founders must therefore “eat depending on the weather.” A healthy business will differentiate into ecosystems, infrastructure, and other directions rather than remain trapped in homogeneous competition. That may reflect the thinking of some Chinese Big Tech companies, but he does not believe the global market should work that way.

15. Commercialization: Launch 2 of 4 Use Cases First, Use the Proprietary Model to Withstand a Price War, and Find the Missing Point in a Nine-Point Product

  • Of 4 use cases, the first 2 are likely to be closely linked: website and landing-page creation, plus B2B SaaS-style business translation involving complex logic such as logins, communities, and commerce back ends. App and agent use cases will follow. The target is not ordinary people in the abstract, but the actual demand side of software—primarily entrepreneurs whose profile overlaps with indie developers. A customization project that once cost tens of thousands of yuan may now cost only a few yuan plus some tokens.
  • Opening an account will come with free tokens, enough in principle to complete 1 to 3 projects; payment begins after that. A proprietary model allows the company to switch models later at lower cost: “At worst, we can fight a cost-based price war.” Buying models from others makes that much harder.
  • 宿文 rates the product as “probably a nine out of 10 purely as a product,” but commercialization and sustained iteration may be the remaining critical pieces. After the public beta, the team can reassess based on user feedback—for example, marking itself down from 9 to 7 and then filling in the remaining 1 to 3 points. “There is not much point in building behind closed doors anymore.”
  • The financing narrative has not changed since day one: build an end-to-end coding product, “lock onto one direction and build it quickly.” In between comes the dirty work. “Look at the moon from far away and it is beautiful; look closely and it is all bricks.”

16. The Real Validation Signal Is Not Verbal Praise but Deployment: Telemetry Behind the Deploy Button

  • The go-to-market plan goes directly at the global market: inbound marketing plus SEO, supplying content across 20 to 30 channels including Product Hunt and Hacker News, which have already been “run into the ground by Chinese and Indian users.” The plan also includes KOL distribution, community building, use cases, and omnichannel marketing.
  • Real feedback “does not necessarily come from what users say; it comes from telemetry,” including metrics such as bounce rate. One critical validation signal is whether generated software can be deployed. The deploy process will be integrated with cloud partners and tracked end to end; the team can observe whether users return to revise requirements, upgrade, and add features. 宿文 compared it to opening an iPhone and still having to activate it: that is when genuine use begins.
  • 宿文 agreed with the host’s warning about “web revenue” or “vibe growth.” Payments from early experimenters may be a fleeting illusion. Many tokens “ultimately do not flow to real users; they have not reached real consumer users,” so “this is still very unstable.”
  • On leaderboard gaming, he said some products are built specifically to strengthen one capability and are “pure leaderboard work.” They reveal little about generalization, much like exam cramming: a high college GPA does not prove strong capabilities will generalize in the real world.

17. Coding Becomes Infra in 3 to 5 Years; Under Deglobalization, Investors Can No Longer Get Firsthand Information

  • The productivity transformation has not arrived yet. For now, “technology creates new demand by increasing supply.” But in 3 to 5 years, coding will become infra in the company’s target use cases: “As long as you dial up compute, the code can be completed.” Large companies will no longer need to ask a central platform team for help; indie developers can cheaply build MCPs and demos, and eventually turn them into commercial products.
  • Low-code and no-code have been addressing the code-supply bottleneck since the 1980s and 1990s, but never fulfilled that historical mission. 宿文 accepts the analogy of “prepackaged code,” while stressing that the architecture of the LLM era—built around model flexibility—is “completely different.”
  • His sharpest closing observation was that investors and the world’s best teams have “lost the opportunity to communicate.” “Which investor today has genuinely spoken with OpenAI’s core team?” In the past, outside sensitive sectors such as defense, investors could still engage with teams globally. Today, many are seeing information that has passed through multiple layers of journalist reposting and may not be consistently following first-hand sources such as Anthropic’s blog or interviews.
  • He has no personal solution: “There is no way around it. We accept the market as it is and adapt to it.”