Pioneers Insight Method Research Author
Universal Medical Intelligence: OpenAI's Plan to Elevate Human Health, with Karan Singhal
Back to Episodes

Universal Medical Intelligence: OpenAI's Plan to Elevate Human Health, with Karan Singhal

Summary

  • OpenAI is using health as a major way to make its mission real and deliver benefits at mass scale, with more than 230 million people using ChatGPT for health and wellness queries weekly. ChatGPT Health is planned to connect medical records, Apple Health, and wearables while remaining free without rate limits for all users; Karan Singhal says ads are not planned “right now,” creating an early version of what Nathan Labenz calls “universal basic intelligence.”

  • Labenz reports that frontier models reached attending-physician territory in his son’s case, while the remaining human edge belonged largely to embodied clinical context. HealthBench Hard rose from 0% for GPT-4o when created to roughly 40% for current OpenAI models, versus the “20 range” for current competitors. During his son’s hospitalization, Labenz found frontier models “step for step with the attending oncologist,” but doctors retained an advantage from directly observing breathing, color, and overall appearance.

  • OpenAI’s health program is being built around expert feedback, evaluation infrastructure, and workflow evidence rather than connected patient-data training. More than 250 physicians—about 260, by Singhal’s estimate—contribute through advisers, continuous Slack-based red teaming, and close research translators; ChatGPT for Healthcare underwent nine testing waves over six months. HealthBench spans 5,000 conversations and about 49,000 criteria, while a randomized study at Kenya’s Penda Health produced statistically significant improvements in diagnosis and treatment outcomes.

  • The bottleneck is shifting from medical knowledge toward context acquisition, multimodality, and product integration. Singhal says text performance is already strong outside some subspecialties, but models work best when given the complete record and still need better access to imaging, voice, longitudinal measurements, and subtle physical signals. The likely architecture combines native multimodal representations with tool use, Python, and specialized models: “the value of data that [people] collect on themselves should increase over time as model intelligence increases.”

  • OpenAI expects clinical adoption to accelerate through 2026, with entrenched workflows—not physician protectionism—the main constraint. ChatGPT for Healthcare adds HIPAA compliance, medical-evidence retrieval, and clinician workflows; it launched with eight leading institutions and generated more inbound demand than the team could handle. Singhal’s goal is for AI-assisted care to become “part of the norm of care” by year-end, though he cautions that healthcare changes over months, not weeks.

  • Privacy is being treated as adoption infrastructure: health data receives added encryption, remains segregated from ordinary ChatGPT activity, and is not used to train foundation models. Singhal does not claim more data lacks value; he argues that lowering users’ “activation energy” through a clear privacy bargain will create greater long-term health impact. The next policy frontier may be consented data sharing for trial matching, N-of-1 treatments, and Labenz’s proposed “AI and the right to try.”

  • Healthcare also serves as OpenAI’s applied alignment laboratory, but Singhal does not claim the hardest oversight problem is solved. Medical models already outperform individual physicians in narrow areas, forcing work on scalable oversight, calibrated uncertainty, expert aggregation, and model character; meanwhile, chain-of-thought has not shown broad drift into “neuralese” as reinforcement learning scales. The upside case is medical “Move 37s”—unexpected diagnoses or treatments that raise the ceiling of health—but the balance between rapidly expanding capability and rarer safety failures remains uncertain.

Deep dive

1. Health as OpenAI’s route to broad AGI benefits

  • Singhal began working on healthcare AI roughly four years ago and moved into it full-time around the end of 2022. His starting conviction was that AGI would probably arrive within his lifetime and that he could improve the outcome in two ways: “work on safety” and “work on benefits,” with healthcare the most obvious benefit.

  • Before ChatGPT, he saw a “capability overhang”: scaled, instruction-tuned language models could do remarkable things, while clinical AI still treated them as speculative. His initial ambition was therefore twofold—convince healthcare that LLMs could work, then develop the safety and reliability machinery needed to make them trustworthy.

  • OpenAI’s health effort began with three goals: universalize access to medical expertise, use healthcare to ground safety and alignment research, and bring society along through partnerships, products, and policymakers. Singhal’s change of mind is revealing: goals that sounded highly ambitious two years earlier now “aren’t ambitious enough.”

2. OpenAI uses collective physician judgment rather than only a static rulebook

  • Singhal divides the program into three phases: laying safety foundations, absorbing rapid adoption, and scaling impact. The adoption phase arrived quickly—health became one of ChatGPT’s fastest-growing use cases, reaching more than 230 million weekly users before the dedicated health products were fully deployed.

  • A first-principles specification is transparent, but Singhal argues that neither he nor Labenz can reliably enumerate correct behavior across medicine’s corner cases. “Do no harm” identifies an important boundary yet “doesn’t tell you what to do in most scenarios,” so OpenAI instead aggregates judgment from roughly 260 physicians.

  • The physician network has three layers: high-level advisers shaping strategy; a close Slack community comparing outputs, red-teaming, identifying blind spots, and testing products; and a smaller group that synthesizes the wider community’s judgment into evaluations, training data, and actionable signals for researchers.

  • ChatGPT for Healthcare illustrates the depth of the process. Physician red-teamers tested it through nine waves over six months, while the broader community surfaced demographic, clinical, and communication failures rather than merely completing isolated preference-ranking tasks.

3. “Do no harm” is a floor, not a complete medical decision policy

  • Labenz’s pushback—worth keeping—is that institutional medicine can become excessively conservative. A physician friend privately told him, “I don’t need a randomized controlled trial to tell you that that makes sense,” yet hospital clinicians often resist reasonable hypotheses or actions because biology is messy and success cannot be guaranteed.

  • Singhal locates the problem on both sides: patients may need to advocate aggressively after clinicians normalize warning signs, while clinicians face exploding evidence, documentation overload, and ordinary human limits. In Labenz’s case, abnormal blood results and initial reassurance showed why merely deferring to the existing system can fail.

  • AI can integrate a patient’s history, present condition, and current medical evidence inside one context. Where a doctor might simplify communication by presenting one treatment path, a model might surface three to five—provided it clearly states that the evidence is limited and distinguishes plausible options from established conclusions.

  • OpenAI’s calibration work therefore asks two separate questions: can a model recognize its own uncertainty, and can it verbalize that uncertainty usefully? Singhal sees this as the mechanism for sharing early evidence without becoming reckless, or becoming so conservative that the model withholds its “best guesses.”

4. Reliability is now measured across repeated reasoning, not next-token confidence

  • Older calibration plots compared a model’s probability for a multiple-choice token such as “A” with how often “A” was correct. Singhal says that method no longer captures the task: health answers are richer than multiple choice, and reasoning models emit thinking tokens before their final response.

  • HealthBench adds detailed clinical rubrics, while “worst-of-N” tests consistency by sampling a model repeatedly and recording its weakest result. Sampling 20 times exposes whether different reasoning paths occasionally produce a materially poorer medical response.

  • For a demanding case, a user could sample ten answers, combine them through an “LM council,” and ask another model to synthesize the result. Singhal expects only a marginal gain over one strong run and says GPT-5 Pro performs something broadly similar under the hood; increasing reasoning effort can deliver a comparable benefit.

  • His practical advice is to favor GPT-5.2 Thinking over GPT-5.2 Instant for health, while default reasoning should now serve most people outside the hardest cases. The worst o3 outputs beat the best GPT-4o outputs; GPT-5 nano and OpenAI’s open-source models now perform similarly to o3, while GPT-5.3-Codex and GPT-5.2 Thinking get better answers with less thinking.

5. Chain-of-thought remains unexpectedly legible as reinforcement learning scales

  • Reasoning models created a useful safety side effect: their thinking tokens often explain—in ordinary English—what they are doing. Researchers can inspect that trace for scheming, undesirable strategies, or discrepancies between a model’s apparent reasoning and its final answer.

  • Labenz worried that denser reasoning might produce an internal dialect or “neuralese.” Singhal reports no large-scale evidence that scaling reinforcement learning has caused a continuous loss of interpretability; there are “weird blips,” but not a robust trend, and he does not assume that favorable result will hold indefinitely.

  • Human-readable thought is not directly guaranteed by training. Singhal’s explanation is that models inherit an English-language prior and use English because it remains the easiest route to a helpful answer; OpenAI’s effort to avoid optimization pressure on chain-of-thought helps preserve this empirically useful property.

6. HealthBench turns medical quality into an unsaturated frontier

  • Medical evaluation moved from exam questions toward realistic interaction: Med-PaLM tested broad health answers, AMIE explored follow-up dialogue, and HealthBench evaluates full conversations with laypeople and professionals. Its 5,000 conversations contain roughly 49,000 distinct evaluation axes spanning facts, uncertainty, escalation, reasoning, and bedside manner.

  • The three HealthBench variants embody three design goals. The full benchmark should be meaningful—“if the number goes up,” real-world health should improve; HealthBench Consensus should be trustworthy, with rubric applicability affirmed by multiple physicians; and HealthBench Hard should remain challenging rather than drifting toward 95–100%.

  • HealthBench Hard was constructed adversarially by selecting high-quality examples on which models across providers performed worst. GPT-4o scored “literally zero” at creation; current OpenAI performance is around 40%, while Singhal places current competitor models broadly in the 20 range, leaving the benchmark far from saturation.

  • The grading layer was itself tested against physicians. Doctors compared HealthBench’s model-based grader with other physicians’ judgments, and the automated grader performed better than the average physician—evidence, Singhal argues, that the benchmark’s detailed scoring is unusually reliable.

7. At the bedside, models reached attending level—but not embodied judgment

  • Labenz’s personal benchmark was a 30-day, high-stakes hospitalization during his son’s cancer treatment. Frontier models stayed “step for step with the attending oncologist on almost everything” and were markedly more knowledgeable than residents, transforming both his confidence in treatment and his own mental health.

  • He began with GPT-5 Pro, then ran questions in duplicate with Gemini 3 and in triplicate with the latest Claude. The AIs disagreed with one another even less than they disagreed with the attending, and only about half a dozen disputes arose across the month, usually over minor decisions such as electrolyte replacement.

  • With hindsight, Labenz scored those disagreements roughly “six to four for the doctors”: about two-thirds of the time, following the clinicians appeared right. That narrow human advantage did not arise from the information he could upload, but from evidence unavailable in the PDFs.

  • The attending could add, in effect, “looking at him right now”—watching breathing, color, and subtle overall presentation. Labenz also hedges the result: his son’s cancer was not especially rare and followed a well-established protocol, so this was emotionally difficult care without necessarily being the frontier of clinical judgment.

8. Better context now matters as much as better text reasoning

  • HealthBench tests whether models escalate urgent cases without flooding health systems through alarmism. Its largest single focus is global health: adjusting reasoning for sex, regional epidemiology, care availability, and facts such as whether tuberculosis is common in the user’s location.

  • Models have also improved at browsing when uncertain, synthesizing current resources, and asking prioritized follow-up questions. Labenz learned to instruct the model to interview him and determine whether he could perform parts of a physical exam; newer systems increasingly recognize and propose that workflow themselves.

  • Singhal says text-based performance is already strong outside certain subspecialties. Health work now feeds every major training stage for OpenAI models, following improvements across o1, o3, GPT-4.1, GPT-5, and subsequent releases; the next limitations are often missing context and modalities rather than raw medical recall.

  • ChatGPT Health attacks that context bottleneck by connecting medical records, wearables, and Apple Health. Singhal’s general advice is that models benefit from more context, although he will not claim that a particular bedside photograph would necessarily have changed Labenz’s case.

9. Medical multimodality will mix tool use with native representations

  • Models have long performed strongly on “crossword puzzles for doctors”—difficult clinical cases where all relevant context, sometimes including multimodal data, is supplied upfront. The harder product problem has been conducting a conversation that elicits the missing observations before attempting diagnosis or treatment guidance.

  • Biomedical data has a long tail of awkward formats: gigapixel pathology images, PET scans, and 3D or 4D MRI data cannot simply be pasted into a chat. One path mirrors clinicians, who inspect selected slices through viewing tools rather than perceiving an entire high-dimensional scan at once.

  • Singhal expects models to manipulate those data through Python, purpose-built libraries, or specialized perception models, while other modalities will be encoded directly into token space as images, audio, and video already are. The eventual medical stack will likely be hybrid, chosen modality by modality.

  • Labenz cites SleepFM, which he describes as combining roughly six to eight modalities measured during one night of sleep to predict multiple diseases. His thesis is that “all the latent spaces will be joined”; Singhal shares the optimism but stresses that timing depends on each modality’s data availability, relevance, and impact.

10. Real-world deployment closes the gap between evaluations and outcomes

  • Personalization produces an enormous testing surface: two users may have different memories, connected tools, instructions, demographics, and communication preferences even when using the same model. Labenz worries that this long tail creates failures such as sycophancy and makes a single “bedside manner” impossible to specify.

  • OpenAI combines pre-launch evaluations with privacy-preserving production monitoring, running classifiers over traffic to detect health and broader safety patterns. Singhal says its work on sensitive mental-health conversations found a strong correspondence between offline evaluation results and patterns observed in real usage.

  • Penda Health supplied a stronger test: in what the episode describes as the first randomized real-world study of an LLM clinical copilot, Kenyan clinicians received an AI safety net that flagged potentially important or incorrect decisions while they documented care in the electronic record.

  • Patients treated by clinicians with the copilot showed statistically significant improvements in diagnosis and treatment outcomes. Singhal’s evaluation ladder now runs from medical exams, through HealthBench’s realistic offline conversations, into prospective clinical studies and retrospective production monitoring.

11. ChatGPT Health prioritizes user trust over connected-data training

  • Labenz personally leans toward sharing: public discussion of his son’s illness connected him with useful expertise, and he sees few non-science-fiction scenarios where disclosure would cause comparable harm. Singhal’s answer is deliberately pluralistic—privacy preferences are personal, so infrastructure must support people who choose differently.

  • Ten years earlier, Singhal argues, many patients could neither access their own records nor control where those records moved. Access has improved, but data remains fragmented across institutions, leaving patients unable to form a complete picture or use it for self-advocacy.

  • Data connected to ChatGPT Health will not train OpenAI’s foundation models. Singhal does not say data has ceased to matter; he says existing models are already strong, and a privacy-first bargain lowers the “activation energy” for users whose fear of training use would otherwise prevent them from receiving value.

  • The product adds health-specific encryption and segregates health information from ordinary ChatGPT activity. Health conversations can benefit from permitted memories or connected context, but regular conversations do not receive the user’s health context in return.

12. The quantified-self payoff rises with model intelligence

  • ChatGPT Health changed Labenz’s view of wearables: he bought a WHOOP because AI could finally process the data without turning quantified-self analysis into another job. His broader personal metric is whether AI lets him spend less time at his desk and more time exercising.

  • Singhal’s call is unusually direct: “the value of data that [people] collect on themselves should increase over time as model intelligence increases.” Better research and product capabilities should extract insights unavailable today, so “if there was ever a time to get a watch,” he says, it is now.

  • Labenz’s daily model comparison puts Claude, Gemini, and ChatGPT on the same laboratory PDFs; “Grok has not cracked that top tier” in his testing. Gemini 3 is most assertive about advocacy and reassurance, ChatGPT is longest and most clinically neutral, and Claude is briefer but more measured than Gemini.

  • Singhal offers no universal winner between thoroughness and digestibility. OpenAI explicitly trains models to infer whether the user is a professional or layperson and adjust jargon and detail—an area he says improved in GPT-5.2 Thinking—while competitors help move medicine’s broader “Overton window” toward acceptance.

13. AI may force a new bargain around experimental treatment and data

  • Labenz imagines mounting pressure when AI identifies an experimental treatment, argues that it is a patient’s best remaining option, and the medical system still denies access. His proposed bargain, “AI and the right to try,” would exchange broader access for systematic capture of each patient’s outcome.

  • That could support N-of-1 decisions—society’s best individualized guess from available evidence—without necessarily abolishing clinical trials. The key is to fold results back into collective medical understanding, allowing humans and models to learn from interventions that otherwise occur invisibly or never occur.

  • Fragmentation harms every participant: patients cannot assemble their histories, clinicians cannot retrieve complete records across systems, and researchers struggle to find eligible participants. Singhal says recruitment challenges cause most clinical trials that fail before conclusion, making consented matching a direct route to faster research.

  • Singhal preserves the counterweight: conventional trial standards are “battle-tested,” and experimental access would depend on active, informed consent. His preferred sequence is privacy-first deployment now, followed by additional consent mechanisms through which patients can share data, find trials, explore treatments, and advance science voluntarily.

14. Health access is being separated from advertising and subscriptions

  • Labenz’s 40-page hospital consent packet—and a later, seemingly duplicative request from the University of Minnesota—shows the friction in today’s public-good data collection. Even a highly motivated participant can fail to return paperwork, implying poor yield from systems built around repeated forms and manual outreach.

  • Singhal draws a bright line around monetization, with an important time hedge: “ads aren’t coming to ChatGPT Health, and we don’t plan for that right now.” The rationale is to separate health impact from incentives that users might reasonably suspect could distort medical recommendations.

  • ChatGPT Health is also planned to be free without rate limits for all users, an access level unavailable elsewhere in the product. Singhal acknowledges that practical limits may emerge but sees no immediate trade requiring ads or data-for-subscription exchanges; Labenz calls the model “universal basic intelligence.”

15. Physician adoption is a workflow problem, not a protectionism problem

  • Consumer adoption has moved faster than clinician adoption, meaning doctors often first encounter medical AI through patients. Labenz’s experience is typical: initial skepticism can evaporate when he shows a provider the model’s actual answer and the provider responds, “Oh, okay, pretty good.”

  • ChatGPT for Healthcare, announced on January 8 after ChatGPT Health, is the institutional counterpart: HIPAA-compliant, purpose-built for clinicians, equipped with medical-guideline retrieval and enterprise writing workflows. It launched with eight leading institutions and produced more inbound interest than Singhal’s team could handle.

  • Singhal reports much less guild protectionism than Labenz expected. Clinicians and health-system executives who have personally used the models often become easy conversations; the constraint is not missing top-down demand but entrenched workflows that cannot change at software-industry speed.

  • Penda’s deployment required active change management—training sessions, peer learning, and help integrating the tool into practice. Singhal nevertheless expects “one of the faster rollouts of software in healthcare history” and hopes AI-assisted care becomes part of the norm by the end of 2026, though not within weeks.

16. Healthcare doubles as a concrete alignment laboratory

  • Singhal joined OpenAI partly because medium- and long-term alignment research felt grounded in toy settings or mathematics. Healthcare supplies immediate stakes, expert feedback, and short-term incentives, turning abstract questions about trustworthy superhuman systems into problems researchers must solve for deployed products.

  • Scalable oversight already appears in miniature because models can outperform physicians on narrow tasks while still requiring physician supervision. When individual experts cannot reliably evaluate every output, training and evaluation must aggregate many experts, improve their ability to critique, and identify failures invisible to any one reviewer.

  • HealthBench, the physician network, and high-compute reinforcement learning are therefore alignment work as well as medical work. The objective is to learn from the collective structure of expert judgment without pretending that one doctor—or one written specification—fully defines safe behavior.

  • Singhal ultimately frames the problem as model persona, character, or “soul.” As models receive more patient context and gain more actions, the decisive question becomes whether they reliably do the right thing; training must surgically preserve useful initiative while avoiding both reckless certainty and reflexive over-conservatism.

17. Scalable oversight is becoming ordinary post-training—not a solved problem

  • Singhal now divides scalable oversight into “rater scaling”—eliciting better judgments from humans and experts, potentially with AI assistance—and “value oversight,” spending enough compute to instill those judgments into a model. The vocabulary has changed even as the original research problem persists.

  • Specs, constitutions, and systematic adherence training are examples of value oversight. Progress increasingly resembles “humdrum post-training” rather than a separate superalignment breakthrough, but Singhal views that grounding positively: system cards show improving adherence and broader generalization of selected personas and safety requirements.

  • Singhal points to a related reason for optimism: critique or discrimination appears easier than producing the correct answer. A model can often monitor itself, especially when given privileged access to chain-of-thought. The discussion also considers whether a less capable model with more inference compute could help supervise a stronger one, but Singhal does not claim that this scaling regime is understood.

  • Singhal does not close the loop: trusted, less-capable models aligning substantially more-capable systems remains incompletely understood. He is encouraged that safety generalization strengthened at larger reasoning scales, yet the net result of expanding capability surfaces and declining per-task failure rates is still “hard to predict.”

18. Medical “Move 37s” shift the ambition from raising the floor to the ceiling

  • Labenz asks what medicine’s equivalent of AlphaGo’s Move 37 would be: a choice no human would make that later proves brilliant. He also proposes training systems to predict a patient’s future health under changed variables, potentially enabling intervention forecasts or in-silico experimentation rather than merely answering questions.

  • Singhal thinks primitive versions may already occur when patients see multiple doctors, then ChatGPT flags the overlooked possibility that leads to a diagnosis. Whether that qualifies as Move 37 is “a matter of taste,” but clearer demonstrations should emerge as models improve their understanding of individual biology and longitudinal health trajectories.

  • Full patient simulation is expensive and not what current models are primarily designed to do. More tractable is predicting an outcome or intervention effect from rich context; Singhal believes pre-training, reasoning-style training, added modalities, and extensions of current paradigms still contain “quite a lot of juice.”

  • OpenAI’s three channels are consumer navigation, AI-assisted health systems, and research that raises medicine’s ceiling. Singhal’s closing thesis is that many biological breakthroughs lacked no physical prerequisite—only sufficient ingenuity—so long-running agents connected to the right data could create an AI biomedical scientist alongside the AI doctor.