Pioneers Insight Method Research Author
E238 | On AI-First Organization Design in the Harness Era: From Trusting People to Trusting AI
Back to Episodes

E238 | On AI-First Organization Design in the Harness Era: From Trusting People to Trusting AI

Summary

  • True AI-first means handing the initiative over productivity from people to AI, then rebuilding workflows, permissions, and the organization around it—not simply giving every role an AI tool. 陈凯 says that if people remain tool users, efficiency gains may top out at 10x; reaching 100x or 1,000x requires AI to lead planning, execution, and cross-team alignment, while people provide signals, architect the system, and review the output. For investors, the dividing line is not AI adoption but whether a company can move from “trusting people” to “trusting AI” backed by guardrails and process.

  • Harness engineering lifts the competitive battleground from models and prompts to a complete system that can run continuously, self-heal, and improve itself. It covers context, tooling, sandbox-host interaction, security, startup time, latency, testing, and the feedback loop; the longer a system runs and the more tools it uses, the more likely it is to encounter hallucinations, context overflow, and capability degradation. Peter’s non-consensus view is how “a system can genuinely come alive from a static state,” with AI—not people—driving iteration.

  • Creo’s internal efficiency benchmark is that a product once requiring about 100 people and 4-5 months can now reach its first deployment with fewer than 10 engineers in roughly 2 weeks. 泓君 quoted Peter’s post saying that at the roughly 25-person company, AI writes 99% of the code: a feature is proposed at 10am, A/B-tested at noon, underperforming parts cut at 3pm, and a new version rewritten by 5pm, versus a traditional cycle of about 6 weeks; the post drew roughly 1.87M views. The investable shift is not simply lower headcount, but that iteration speed turns the product roadmap from a scarce resource into inventory.

  • Whether AI coding can scale depends on whether the quality loop can turn bugs from manual firefighting into automated discovery, triage, and repair. Creo’s agent-driven CI/CD and bug triage can detect an issue in 1-2 minutes, assign it in seconds, and investigate, fix, and redeploy within 1-2 hours, versus potentially a week before; more than 50% of issues are reportedly handled through auto-fixing, with engineers only needing to approve low-risk changes. Sensitive changes involving security or agent behavior still receive deep human review, showing that scarce human value is shifting toward architecture and high-risk judgment.

  • Once development speed exceeds the market’s ability to absorb new products, the bottleneck moves from engineering to go-to-market, demand judgment, and evaluation. Creo says its technical team is already 3-4 months, and in some cases 4-5 months, ahead of marketing: when the market wants apples, it takes apples from the “basket”; when it wants bananas, it takes bananas. Articles, videos, and marketing assets lack code’s clear correctness metrics and are still produced as many agent outputs for humans to judge. The investment implication is that exploding code supply could make demand sensing, content evaluation, and the feedback-data chain the next constraints.

  • Falling development costs are re-rating product managers, engineers, and designers—not simply eliminating one of the roles. Creo has distributed product-manager responsibilities across engineers and engineering managers, with the broader team likely to share the function going forward. Peter says the truly scarce talent is the hybrid operator who can take an idea directly into the product within 1-2 hours, because once a handoff is required, “the cost of alignment is far greater than the cost of implementation.” Senior engineers who embrace AI remain irreplaceable, but a company may need only 1-2 of them rather than 10x or even 50x as many as before.

  • Agents may become both software operators and first-line consumers of content, advertising, and procurement information. When choosing task-management software, the team is shifting its focus from human-facing dashboards to agent-ready MCP and APIs; when Clark sourced SOC 2 and ISO compliance services, he first had an agent research and screen the options. Creo therefore advocates moving from one shared general agent to a dedicated agent for each company, with continuous self-healing and self-improvement, initially targeting SMBs with about 30 or fewer employees and relatively light legacy and compliance burdens.

  • The guests are broadly optimistic about the long-term outcome but clearly cautious about the transition. Human value is distilled into defining the need, architecting the system, deciding whether something still has value, and reviewing the final result; Peter also stresses that even today, AI cannot replace people in anything. The real risks sit in the transition: organizations unwilling to relinquish control, permissions and privacy boundaries that remain unresolved, senior professionals’ legacy specialties being re-rated, and large numbers of roles having to be redefined amid pain and noise.

Deep dive

1. Harness Expands Model Optimization into a Complete Dynamic System

  • Peter draws a clear distinction between the three stages: prompt engineering optimizes prompts, context engineering fills in context, and both mainly address human-LLM interaction; harness engineering covers the infrastructure surrounding the model.

  • Unlike static optimization for a single vertical task, a harness must handle a general system: how tooling is connected, how the sandbox and host interact securely, how long the sandbox takes to start, how high latency is, and how the system recovers after failure.

  • 泓君 summarized it as “how to push an LLM to the outer limit of its best use.” Her contrast: one agent completed the work of 3 SEO employees overnight, while another content pipeline ran for 2 days before anyone discovered that “everything was garbage.” The difference lies not only in the model, but in whether the surrounding system can continuously correct itself.

2. The Real Non-Consensus View Is Making the System Come Alive

  • Peter says the market still tends to understand harnesses statically: build a set of constraints and tools to get more out of an LLM. Creo’s definition is instead: “How can a system genuinely come alive from a static state?”

  • A dynamic harness continuously receives marketing, product, user, and infrastructure signals, then lets AI lead the iteration. The human role is not to edit each output, but to “feed all kinds of signals to AI.”

  • An agent is not complete when it is created in a one-shot process. How users follow up, and how the system self-heals and improves itself based on results, determine whether it can carry real work over time.

  • Peter views long-running tasks as scaling during inference: provide more context, more tooling, and more time to think. The trade-off is a simultaneous rise in the probability of hallucinations, context overflow, and model capability degradation, making the harness itself a complex engineering problem.

3. AI-First Is Not Adding a Copilot to the Old Workflow

  • 陈凯’s core judgment is: “It’s not about using AI tools on top of existing processes; it’s about rebuilding your workflows and organizational structure around AI’s capabilities.”

  • Creo once had engineers use AI to write code, product managers use AI to write PRDs, and designers use AI to create images. Overall efficiency did not improve meaningfully. Once everyone’s individual speed increased, the remote team’s working rhythms diverged further and alignment costs rose instead.

  • 陈凯’s ceiling is aggressive: as long as people remain the users of productivity tools, gains may top out at 10x. To reach 100x or 1,000x, “AI should lead all productivity,” with people shifting from hands-on workers to reviewers of results, system collaborators, and judges of value.

4. Organizational Transformation Changes Trust Before It Changes Code

  • Peter says the first obstacle to AI-first is a people problem: whether the team can accept a new way of working. In the past, it might take months to prove that a new architecture was better; with AI assistance, the team can restructure the front end, back end, and infrastructure in 1-2 weeks, then persuade people through deployment frequency, reliability, and end results.

  • 陈凯 attributes organizational friction to a lack of trust. Previously, GTM and engineering had to communicate, reach consensus, and then move forward; now AI can synchronize upcoming features and their structure directly with marketing, reducing interpersonal alignment.

  • 泓君 asks the practical question: how would AI know that engineering can really finish tomorrow? Peter’s answer is not to have AI guess the schedule, but to reduce reliance on detailed planning and focus instead on whether a feature improves top-line metrics and generates real usage data.

  • Once the data chain is in place, the agent can use launch performance to judge whether a feature works and whether to rule it out or fall back. In other words, alignment shifts from human promises to operating signals that can feed back into the system.

5. AI-Driven CI/CD Turns Quality Control into a Closed Loop

  • Peter does not deny that AI coding creates bugs: “Whether code is written by AI or by people, it will produce bugs.” The goal of a harness is not a static system that never fails, but one that continuously discovers problems and improves.

  • The first line of defense is integration and regression testing, preventing obvious problems from reaching production. Traditional CI/CD is driven largely by rules and unit tests; Creo adds AI-driven integration, unit, and end-to-end testing, with Peter citing Playwright as an example of full-path checks.

  • The second line of defense comes after launch: the system monitors logs, errors, and incidents, feeds those signals back to AI, assesses code quality, identifies corner cases or risk conditions, and then triggers bug triage.

  • Multiple agents can work in parallel on the front end, back end, and core agent system. Peter’s cycle is 1-2 minutes to find a bug, seconds to assign it, and another 1-2 hours for engineers to use an agent to investigate and propose a solution, versus potentially a week under the traditional process.

6. Auto-Fixing Raises Speed but Does Not Eliminate the Architect

  • Creo used to maintain both a feature wish list and a bug list, with marketing, product, and engineering constantly debating whether to build features or fix problems. The team now says both lists have disappeared: bugs are fixed promptly, while feature supply far exceeds current demand.

  • Auto-fixing grades risk by code directory. Issues in low-risk files generate an AI-submitted PR, which an engineer can approve with a simple review before it reaches production. More than 50% of issues are currently handled this way.

  • Changes involving security or agent behavior still require deeper review by the relevant people. Peter’s definition of behavior includes not only outputs, but also cost, latency, hallucinations, tool authentication, and any system interaction that could prevent a user task from being completed.

  • Asked whether engineers still need to relearn the entire codebase, Peter divides the engineering team into architects and operators. AI can provide the solution, but architects still decide on security, latency, and the sandbox-host boundary. A system that once required 10-20 people to build might now be completed by 1 architect in a week.

7. When Errors Happen, Fix the System—not Just the Output

  • 陈凯 says to “treat AI as a system, not as an intelligence.” When an error occurs, the priority should be to find the system vulnerability rather than correct that one output.

  • He also rejects the idea of a harness as a “static, fixed shackle.” A better analogy is raising a child: allow it to keep growing within a set of rules. The platform should not provide a one-time answer, but act like a “trainer” that teaches users how to make the agent perform better after an error is found.

  • Using editing as an example, 泓君 notes that one major mistake can force an editor to reread the entire piece, with costs approaching a rewrite. 陈凯’s answer is to first ask why the system allowed that class of error to occur, rather than spend all the effort patching a single instance.

  • His follow-up question is more radical: is an error in human eyes still an error in the eyes of the final consumer? “Rethink the question of whether it is still a problem”—the first step is to establish who will consume the result in the future.

8. Agents May Become the First-Layer Consumers of Content and Software

  • 陈凯 expects articles, images, and videos to be consumed first by agents in the future. A marketing asset that does not fit human aesthetics might generate better data feedback when read or screened by agents. Harnesses therefore need to optimize for the preferences of the actual consumer.

  • 泓君 remains cautious: agent purchasing and reading are not surprising, but she did not expect them to arrive this quickly. Clark offers a firsthand example: when looking for SOC 2 and ISO compliance services, he first had an agent research and screen the options before reviewing the candidates himself. Second- and third-layer judgments may also be taken over in time.

  • Peter sees the same trend in SaaS interaction. Products such as Asana and Linear were built around dashboards for human users; the team now cares more about whether agents can read, sort, and process tasks, making MCP and API quality more important than before.

9. The Market’s Obstacle Is Often Not Knowing the New Workflow Exists

  • Peter believes the main challenge for most agent companies is not user resistance, but that “the market doesn’t know this way of working exists,” or how to make an agent work effectively.

  • Creo is therefore trying to reduce configuration requirements. Ordinary users do not need to understand sandbox, tooling, security, or long-running infrastructure; they only need to understand their own task, while the underlying harness is provided as a cloud service.

  • This also answers the objection that every vertical requires an expert. The platform does not claim to possess all industry knowledge; it packages the general problems of architecture, security, tool connections, and agent maintenance. Domain judgment remains with the user.

10. Creo’s Transformation Turned on the Shift from Assisting People to AI Leadership

  • 陈凯 recalls that in the first half of 2025, the company still viewed AI as an assistant and people as the leaders. By the second half, it found that efficiency was far below expectations, largely because the “user of the productivity tool” had not actually shifted from people to AI.

  • The change did not happen overnight. Marketing and engineering spent 1-2 months repeatedly discussing a new way to collaborate; around August-September 2025, the team realized it had to rebuild, first aligning on the mindset, and only in January 2026, before the Spring Festival, did it formally change the code and processes.

  • Peter adds that the transformation was constrained by model capability. Foundation models, agent architecture, and infrastructure all improved sharply within a year. Asking AI to lead development a year earlier “was technically impossible”; by the time of the rebuild, both the speed and the quality were there.

  • The actual rebuild took about 2 weeks and covered the front end, back end, architecture, and infrastructure. The product shown on the program came from this reconstruction, rather than from continuing to layer AI features onto the old system.

11. Planning Rose from a Failing Grade to 90, Leaving People to Find Faults

  • Peter illustrates the capability shift with scores: a year ago, AI planning was roughly 50 out of 100, requiring people to directly modify the plan and architecture; now a first draft can reach 90, with people only needing to criticize it before asking AI to revise it.

  • Under the new architecture, he says, “I haven’t written a single line of code or changed a single line of text in the plan.” The interaction consists of probing for security, latency, and architecture flaws, asking AI to reference popular open-source agent frameworks, and then generating the final version.

  • When 泓君 asks whether AI is already more capable than he is, Peter draws a distinction: AI coding is certainly stronger than his current ability—“I haven’t written a line of code in 2026”—but the architect’s value remains in finding security and latency flaws in the plan.

  • A correction can be codified as a skill. For example, the security principle governing the sandbox-host boundary can be written into a skill; the next team member only needs to ask AI to follow that principle rather than explain the rules again. Experience is thereby converted from individual memory into an organizational harness.

12. After R&D Efficiency Surged, Product Supply Moved Months Ahead of the Market

  • Peter estimates that under the non-AI-led approach of a year ago, building the current Creo product would have required at least a team of about 100 people and 4-5 months. Today the company has about 25 people, fewer than 10 in engineering, and can reach its first deployment in about 2 weeks.

  • 泓君 repeated details from Peter’s post, which drew about 1.87M views: write the feature at 10am, run an A/B test at noon, cut part of it based on the data at 3pm, and rewrite a better version by 5pm. The traditional development cycle might take 6 weeks. The post also said AI wrote 99% of the code.

  • In the software era, sales typically led the product by 4-5 months. Creo says the relationship has reversed: engineering is 3-4 months, and sometimes 4-5 months, ahead of marketing, completing capabilities that the market team does not yet know about. Operations are no longer organized around a scarce roadmap, but around which capabilities are worth taking to market.

13. Go-to-Market Is Becoming a Harder Bottleneck to Evaluate Automatically

  • Clark uses a “basket” and Doraemon’s “magic pocket” to describe the new state: if the market needs apples today, the company takes apples from the existing product library; if it needs bananas tomorrow, it does not have to wait for R&D to reorder the roadmap.

  • But GTM is harder to harness than coding. Code has relatively bounded evaluation and clear pass criteria; the value of an article, video, or image is more subjective across different people or agents. Converting those judgments into signals the system can use remains one of the biggest challenges.

  • Clark’s candid boundary is that the company has not yet “let agents make decisions 100% of the time.” For now, it generates many agent outputs and has people judge their quality; even when many features are already available internally, the company will not launch them if it believes the market is not ready.

  • A more advanced internal practice is to give agents broad read-write permissions. In the past, if Clark wanted to query a user behavior pattern, he had to find data or ask an engineer to build a table; now an agent can answer in about 3 seconds. But if it reads the wrong data or “starts going rogue,” decisions become contaminated, so permission controls are not yet ready for the market.

14. Dedicated Agents and Organizational Redesign Are Creo’s Product Bet

  • Peter describes the previous generation of general or super agents this way: the platform provides one unique agent, and all users access the same assistant. Creo wants users to create agents of their own that understand specific workflows and continuously improve and self-heal.

  • His example is a weekly campaign. Instead of directing a general agent from scratch every week, the user creates a campaign agent; the platform harnesses its performance, cost, and output in the background, making the workflow more stable with each run.

  • The target customers are primarily SMBs, which 陈凯 describes more specifically as small and midsize teams with about 30 or fewer people. Both technology and traditional companies can transform; the key is not headcount, but whether compliance, legacy databases, and existing processes are light enough.

  • For customers, the entry point is not to copy Creo’s organizational structure immediately, but to use a productized harness first and then gradually consider organizational change. 陈凯 warns that if a founder cannot accept a wholesale redesign of the database, interactions, and product architecture, simply adding an AI feature to legacy SaaS is unlikely to produce a real transformation.

15. As Role Boundaries Disappear, Value Concentrates in Architecture, Execution, and Judgment

  • Creo’s first organizational change is the object of trust. Organizations once trusted people; now they need guardrails and mechanisms that make AI’s planning, decisions, and execution trustworthy to humans.

  • The product-manager function has not simply disappeared; it has been distributed across engineers and engineering managers. 陈凯 sees PMs as the point where market and R&D conflicts often concentrate. If a standalone product-manager role disappears while the system maintains trust, alignment costs may actually fall. In the future, the whole team may “play product manager.”

  • Peter sees the main trend as the rise of hybrid talent. Engineers need product and marketing sense, while product managers and UX/UI designers need implementation capability. An idea is most valuable when it can enter the product within 1-2 hours; if it must be handed to another engineer, “the cost of alignment is far greater than the cost of implementation.”

  • Junior engineers are often more willing to expand their scope from coding into design, launch analysis, and impact assessment; senior engineers may be constrained by 10 or 20 years of specialization. Peter does not conclude that senior talent is replaceable: the scarcest people remain senior professionals who embrace AI and possess architectural and product judgment, but a company may need only 1-2 of them rather than 10x or even 50x as many as before.

16. What People Ultimately Retain Is the Power to Define Value and Review Results

  • Peter defines the core human capability of the future as system architecture: moving from implementing features to architecting and maintaining AI systems. The same applies to GTM, which must build marketing systems that can run autonomously rather than manually produce individual pieces of content.

  • 陈凯 says that as long as technology continues to serve people, people will define the direction of demand and review whether the results serve their interests. Clark compresses it further: “The value of people in the future is judging whether anything still has value.”

  • That judgment still has a clear boundary. Peter explicitly says that even at the current stage, AI cannot replace people in anything; it can, however, already lead many tasks, while humans must still architect a stable system. As agents communicate with other agents or people, new ethical questions will emerge around who can view the content and whether privacy has been violated.

  • All three guests are optimistic, though to different degrees. 陈凯 believes AI may help people separate work and life more effectively; Peter is optimistic as an entrepreneur, comparing the transition with the Industrial Revolution and arguing that people will find new directions after jobs are displaced; Clark is “very cautiously optimistic,” acknowledging the pain and noise of the transition while believing it will ultimately free more time and human potential.