Pioneers Insight Method Research Author
Babysitting the Machine: Glean's Rebecca Hinds on the Hidden Human Labor of AI at Work
Back to Episodes

Babysitting the Machine: Glean's Rebecca Hinds on the Hidden Human Labor of AI at Work

Summary

  • Workplace AI has crossed the adoption chasm without crossing the organizational-performance chasm: 87% of surveyed workers use it, 73% feel more productive, and reported savings average 13 hours a week, yet only 13% see their organization performing significantly better. Hinds cautions that the self-reported savings range from 10 to 14 hours, with about 11 attributed to fully automated output. The investor-relevant gap is no longer access to AI; it is converting local speed into measurable enterprise outcomes.

  • “Bot-sitting” consumes 6.4 hours a week—roughly half the headline savings—because employees must feed context, debug probabilistic failures, inspect outputs, and manually connect tools. About 36% of AI sessions fail badly enough to require substantial rework or a restart, with context feeding and opaque debugging carrying the highest “exhaustion multiplier.” The bottleneck is increasingly the employee as “the integration layer,” not raw model capability.

  • Unrewarded bot-sitting can culminate in “bot-shitting”: shipping AI-generated work whose quality the sender cannot explain or defend. The episode cites 69% admitting to some form of the behavior, while Hinds gives a narrower 40–41% for employees shipping work they could not explain if challenged. What looks like productivity can therefore be “polished nonsense”—or a coordination loop in which one worker turns a bullet into five pages and another compresses it back into a bullet.

  • The enterprise-AI control point may be context and orchestration rather than any single model. Glean’s thesis is that a shared graph connecting mission, goals, projects, tasks, people, documents, and technology can reduce bot-sitting, distinguish authoritative recent information, and route tasks among models or agents. Hinds sees no plausible single-model future because models “leapfrog each other left and center”; enterprises need choice without losing organizational context.

  • Automating the work employees value most can turn AI adoption into a retention problem even when the economics look compelling. Fearful workers may automate familiar, visible tasks to appear AI-native—including customer relationships that gave their jobs meaning—and heavier bot-sitting and bot-shitting correlate with job-seeking, though Hinds stresses that causation is unknown. Her governing rule: “because the technology can do something” does not mean it should, especially when a cited study found 41% of Y Combinator AI startups automate activities people would prefer to keep human.

  • Successful AI culture requires transparent strategy, psychological safety, and incentives for collective value—not token counts, clicks, or theatrical head-count targets. Rob Cross’s research, as Hinds recounts it, finds high-performing organizations up to 5.5 times more likely to measure and reward effective collaboration; promising experiments reward co-creation, peer feedback, and improvement rather than raw output. Mission matters because it can replace some of the coordination formerly supplied by hierarchy: “mission as boss.”

  • The likely enterprise end state is smaller teams, fewer coordination-heavy managers, rebundled roles, and expensive multi-model systems—not an immediate collapse in AI costs. Hinds hears of annual token budgets exhausted in a month, while one large healthcare organization found 70% task overlap across roles. AI-native companies have an advantage, but her warning to incumbents is explicit: do not “copy and paste” an AI-native operating model onto a legacy organization without identifying which inherited capabilities still create value.

Deep dive

1. AI adoption is a human change measured from two angles

  • Hinds’s starting point is change management: “This is a human change just as it is a technology change.” Calling the system “artificial intelligence” puts it psychologically in tension with human intelligence; even effective tools will be resisted or used symbolically unless employees can understand them as amplifying their abilities.

  • The Work AI Index surveyed 6,000 knowledge workers—3,000 in the US and 1,500 each in the UK and Australia—during December 2025 and January 2026. Eight founding members of Glean’s Work AI Institute shaped questions spanning psychology, technology, digital transformation, and organizational design.

  • Anonymous aggregated Glean telemetry supplied an objective counterweight to inherently biased self-reports. Hinds says adoption has powerful network effects: use by managers, teammates, and cross-functional partners matters, so effective change combines policy and visible executive use with bottom-up “AI influencers or champions,” rather than merely “telling employees to use AI or else.”

2. Glean treats organizational context as the missing infrastructure

  • Glean began with enterprise search before mainstream generative AI, addressing how employees find relevant internal information. Its current work-AI platform combines a personalized assistant with agents that automate cross-functional workflows while understanding both the individual’s job and the organization around it.

  • The “bread and butter” is context: a data model connecting work across the enterprise so answers are not generic. Hinds’s desired progression is from reactive assistance toward predictive and proactive AI that identifies what matters in the current workday and recommends priorities without requiring the employee to reconstruct the background.

  • Context must also encode recency and authority, not merely retrieve documents. Employees want different tools for different jobs, and models “leapfrog each other left and center”; without a common contextual layer, that choice produces AI and agent sprawl, disconnected outputs, and more human integration work.

3. Thirteen saved hours produce little visible enterprise transformation

  • The headline contradiction is stark: 87% use AI, 73% say it makes them more productive, and average reported savings reach 13 hours a week—about one-third of a conventional workweek. Only 13%, however, say their organization performs significantly better because of the technology.

  • Hinds keeps the measurement hedged. Depending on how respondents were asked, reported savings ranged from 10 to 14 hours; work they considered fully automated represented roughly 11 hours. These are perceptions, not directly observed efficiency gains, and the organizational-performance item was a Likert-scale question.

  • Nathan’s optimistic interpretation was that this may be a decent start: the survey captured one moment during rapid model improvement, and the downside might be limited even if only 13% report major gains. Hinds confirmed responses across the scale but did not provide a share claiming that AI made their organizations significantly worse.

  • Hinds also resists demanding transformation too quickly. In the best cases, she says executives are seeing as many as 80% of AI initiatives fail because experimentation requires failure; the concern is that shadow use, undisclosed output, and bot-shitting “are not going to get better if all else remains similar.”

4. Bot-sitting is the hidden labor inside AI productivity

  • Hinds defines bot-sitting as feeding AI context, overseeing its work, debugging failures, and cleaning up afterward. Workers report spending about 6.4 hours weekly on it, nearly half the headline gain; much of that labor is tedious, untracked, unrewarded, and generated by fragmented tools rather than employee shortcomings.

  • The report divides AI time into bot-sitting, using AI interactively to advance real work, and learning or building agents. Roughly 36% of sessions fail: the worker must either start over or perform substantial rework. Better first-pass context could redirect those hours into production or capability-building.

  • Feeding context and debugging create the strongest “exhaustion multiplier.” The first feels like supplying information the system should already know; the second is frustrating because probabilistic systems rarely reveal which component broke or why a small prompt change worked, leaving the employee to probe a black box.

  • Nathan’s pushback—worth keeping—is that many enthusiasts gladly trade old manual work for AI supervision. At an AI event, the median attendee estimated that two unaided people would be needed to replace one person using AI. Hinds’s answer: curious experts are outliers; many workers are “too exhausted to be curious right now.”

5. Bot-shitting converts local speed into organizational stasis

  • Hinds’s proposed cycle begins with adoption pressure, which increases bot-sitting without increasing recognition or incentives. Exhausted employees reach “good enough”—the research term is “satisficing”—and treat a plausible-looking response as permission to ship, converting hidden labor into unowned output.

  • The episode cites 69% admitting to some form of bot-shitting; Hinds separately says 40–41% ship AI work they could not explain if asked. The broader category includes shadow AI and other unaccountable use, while its most visible artifact is “polished nonsense”: finished-looking work with little substance underneath.

  • Nathan’s own near-miss showed the mechanism. His preparation agent searched his calendar, Drive, prior interview outlines, and the web, but missed the new report delivered by email; 80% or more of its proposed conversation was therefore beside the point. He caught the error, supplied the report, then still internalized the material before interviewing.

  • “Coordination neglect” explains how individual savings disappear: one employee expands a bullet into a five-page AI report, and the recipient compresses it back to one bullet. Both look faster, but the company gets a “hamster wheel of AI slop.” Employees may also hide saved hours because disclosure could simply earn them six more hours of work.

6. Automation can remove the work that made a job worth doing

  • One paradox is that workers most afraid of replacement may adopt AI most aggressively. Lacking confidence or organizational support, they want to appear AI-native; the most visible material to automate is usually the work they know best, which may also be the work that gives them meaning.

  • Customer service carries the argument. A representative may have spent years or decades cultivating human relationships, only to be reassigned from talking with customers to configuring and supervising agents: “I didn’t sign up for this.” The technology has not merely removed effort; it has substituted an unwanted occupation.

  • Nathan’s counterargument is economic and customer-centered. Intercom’s Fin can respond within minutes where human back-and-forth might take 30 minutes or more, and leaders may confront a hypothetical 90% saving alongside better responsiveness. Hinds concedes that firms should sometimes automate meaningful work—but, in the best case, replace it with work employees find equally or more meaningful.

  • Hinds invokes the “IKEA effect”: doing difficult, friction-filled work builds ownership, judgment, purpose, and pride, which are performance drivers rather than decorative benefits. She cites a Stanford study finding that 41% of Y Combinator AI startups automate tasks people would prefer to keep human. Capability alone cannot determine the division of labor.

7. An enterprise graph could allocate humans and agents dynamically

  • Hinds defines the enterprise graph expansively: mission, goals, projects, tasks, people, documents, and technology—not merely a map of files and reporting lines. That context could let AI evaluate work against both business objectives and the way the organization actually operates.

  • In customer service, the graph could inspect prior interactions and infer whether a request calls for fast automation, a human-in-the-loop process, or relationship-building by a person. Complexity and customer preference become inputs rather than applying one automation percentage to every conversation.

  • The same machinery could allocate work using employee expertise, development goals, passions, and bandwidth alongside company priorities. Instead of staffing from a static org chart, AI could search a base of 1,000, 2,000, or 10,000 employees and recommend a project team through a calculus no manager could perform manually.

  • Respondents in the 13% who say significant organizational productivity gains are occurring work in organizations that measure more than productivity and disproportionately put the resulting data in employees’ hands. Hinds saw similar benefits with collaboration technology and hybrid work: transparency helps employees understand the system, while a queryable graph could make organizational state broadly visible rather than reserving it for management.

8. Detection must diagnose why people cross the guardrails

  • Nathan’s provocative reading of the 69% figure is that AI may already work remarkably well: if two-thirds of workers pass along some unowned AI output and “the wheels aren’t falling off entirely,” many activities may be more automatable than leaders realize. He asks whether a detector—likely Pangram Labs—plus a quality score could reveal where.

  • Hinds envisions a world where tools report response uncertainty and estimate pure-AI versus human-AI generation. Enterprise context could improve that judgment by comparing an artifact with an individual’s normal style—“Rebecca’s default writing style”—instead of relying on generic linguistic signatures that make present detectors unreliable.

  • Technology is only part of the control system. Drawing on Amy Edmondson’s work, Hinds argues that psychological safety should let employees say, “This is bot-shitting,” including their own contribution and its cause. Identifying the output without understanding the incentive or tool failure behind it leaves the cycle intact.

  • Shadow AI illustrates the danger of punishment alone. Employees using an unsanctioned tool introduce real risk, but Hinds says they are often high performers coloring outside the lines because the approved stack fails them and they can see the productivity upside. The gold standard is to make “the safe path the more efficient path.”

9. Heavy AI use can signal both flight risk and rising value

  • Greater bot-sitting and bot-shitting are correlated with active job-seeking, but Hinds is explicit that the survey cannot establish causation. The two behaviors may point to different mechanisms, so leaders should not treat every intensive AI user as either a star or a disengaged employee.

  • For bot-sitters, Polly Annardi’s work on “digital exhaustion” offers one explanation: “The digital employee experience is increasingly the employee experience.” Constantly supplying missing context undermines confidence in an employer loudly proclaiming AI transformation; workers may leave for an organization whose tools make that strategy credible.

  • Bot-shitting may reflect a later stage of disengagement: employees no longer feel ownership of what they send. Aruna, a report contributor from Berkeley, offers another hypothesis—the employee may have become so capable with AI that their market value is now higher outside the organization than inside it.

  • Hinds sees little meaningful compensation for elite enterprise AI collaboration today, though she thinks there should be. Rob Cross’s research finds high-performing organizations up to 5.5 times more likely to measure and reward collaboration; promising hackathons and agentathons recognize improvement, before-and-after prompts, co-creation, and peer feedback—not only the largest headline impact.

10. AI transformation fails when leadership becomes theater

  • “Employees can call bullshit from a mile away,” in Hinds’s framing: a collaborative talk track cannot coexist credibly with unexplained cuts. She has heard executives begin with a 15% head-count target and debate whether it should be 14% or 16%, a performative precision detached from the organization’s actual work.

  • Flattening, layoffs, and hierarchy reduction are organizational-design changes, not generic AI recipes. Firms sometimes remove customer-service staff, discover those people carried indispensable long-term relationships, and bring them back. Moving before mapping the work produces regret that a richer enterprise graph might prevent.

  • Mission becomes more valuable as hierarchy recedes because hierarchy once told employees what to do under uncertainty. A believed mission can become the replacement decision rule—provided people understand how their own work ladders into it. Without that connection, purpose cannot prevent buck-passing or symbolic adoption.

  • Nathan offered Elon Musk’s reduction of Twitter from roughly 7,500 employees to perhaps 1,000–1,500 at its low as a case that might influence executives. Hinds declined to validate that case specifically: Jensen Huang’s “mission as boss” culture and lack of one-on-ones with direct reports at NVIDIA may work there, but copying a celebrated leader’s visible practice into another culture is dangerous.

11. AI works best as a teammate whose mistakes remain human-owned

  • The teammate metaphor gives employees a usable mental model. Unlike a hammer or calculator, a teammate is not transactional or expected to produce perfection immediately; value emerges through interaction. At Glean, Hinds consults her assistant throughout the day to move work forward rather than treating each query as isolated.

  • The metaphor has a limit: AI is not a peer employee to whom blame can be delegated. Citing Leonardi’s research, Hinds notes that recipients still blame the human when an AI assistant errs. Agentic capability does not transfer responsibility away from the person deploying the output.

  • While describing her time joining Glean, Hinds says she uses the assistant to recover institutional context: why a feature launched, who customers are, what the roadmap contains, and how different executives prefer to consume information. On joining, she asked what successful employees at Glean do differently and adjusted accordingly.

  • The answer emphasized team performance, long-term thinking, and willingness to raise a “weird, wacky idea.” Her assistant now learns her priorities through memory, lets her select models by task, flags unanswered emails, and surfaces action items left in documents—proactive context that removes the file-shoveling form of bot-sitting.

12. Smaller teams will not mean cheaper AI in the near term

  • Hinds rejects an easy near-term story of falling AI costs. Executives sometimes consume an annual token budget in one month; the likely architecture is therefore multi-model and multi-tool, with AI routing each task according to efficiency, complexity, and cost rather than using the most expensive capability indiscriminately.

  • She does expect smaller teams and perhaps fewer managers because AI can reduce work’s “massive coordination tax.” Specialists may become broader generalists, while roles formerly fragmented across two, three, or four positions can be rebundled once AI exposes work across silos.

  • One CHRO at a very large healthcare organization used AI to map tasks and found 70% overlap across roles. That creates a concrete design problem for the enterprise graph: which activities belong in a human role, which belong with an agent, and what agent-to-human ratio fits this organization?

  • Hinds remains conditional about the outcome: AI is not inherently good or bad, and benefits depend on intentional treatment of the human system. AI-native companies have less organizational baggage, but legacy firms must preserve what already works and evolve selectively—never “copy and paste” an AI-native business model onto themselves.

13. Research and meetings still require grounded human judgment

  • The Work AI Index is intended to become a pulse survey, repeated roughly every six months to track bot-sitting, changing roles, and organizational outcomes longitudinally. AI can accelerate analysis, but Hinds and her co-authors kept parts of the report’s narrative and storytelling “100% human.”

  • Her ethnographic training still matters: embedding inside an organization for months or years reveals changes that surveys, interviews, and telemetry cannot. Glean is already experimenting with AI inside meetings, while her ideal extreme case would be a legacy organization radically rewiring its org chart and allowing researchers to observe the consequences on the ground.

  • Meetings show both sides of AI. Systems can assess meeting health, detect executives dominating airtime, distinguish creativity from coordination, and automatically remove sessions lacking sound design, an agenda, or accepted participants. Used this way, AI protects expensive synchronous time.

  • Sending digital twins or note-taking bots instead of attending can be cognitive offloading disguised as modernization. If a bot can substitute entirely, the meeting may never have been needed; sending one also signals that the organizer does not value colleagues’ time. The prerequisite remains decidedly nontechnical: know what genuinely deserves to be a meeting.