Skip to main content

The Hard Part Isn’t Vibe Coding—It’s Getting AI Safely Into Healthcare Workflows

Dr. Rifat Atun speaks at the Pressure Points webinar.
Dr. Rifat Atun speaks at the Pressure Points webinar.

Panelists at a Harvard Chan Studio event discuss why the real challenge in health AI begins after the prototype works

Artificial intelligence is changing what it means to build in healthcare. That was the focus of Harvard T.H. Chan School of Public Health’s recent Pressure Points event, Engineering AI for the Future of Health Care, co-hosted by The Studio and the Advanced Learning Academy (ALA). Moderated by Rifat Atun, Dean of Education and Innovation at Harvard Chan School, it featured Aman Bhandari (SCAN – Senior Care Action Network), John Brownstein (Boston Children’s Hospital), Heather Mattie (Harvard Chan School), and Trishan Panch (LUNRStudio and Lumin Health), who co-directs ALA’s Responsible AI for Health Care, Innovation with AI in Health Care, and Beyond Vibe Coding: Building AI Solutions to Transform Health Care courses.

Atun opened by naming the shift now facing healthcare leaders: AI is no longer only a tool for automation, synthesis, or analysis. Increasingly, it is “a way to build.” Agentic AI and vibe coding, he noted, are letting clinicians, public health leaders, and researchers generate prototypes and build tools faster than before—raising new questions about what leaders should build, buy, and scale.

From Coding Syntax to Building by Intent

Panch described vibe coding as a change in who can participate in software creation. Rather than writing exact code syntax, users can describe what they want and let AI generate the code behind it—what he and Brownstein call natural-language programming: using AI agents to pursue multi-step goals through ordinary language. Because AI can now write code at a professional level, Panch said, “many more people can now come into the fold of those that can create software.” For healthcare, that matters because the people closest to clinical, operational, and public health pain points may be best positioned to identify what needs to be built.

But easier building does not remove the need for discipline. A prototype that works in a demo is not the same as a tool that can be trusted in care delivery, research, operations, or public health practice.

The Deployment Gap in Healthcare AI

At Boston Children’s Hospital, Brownstein said, AI tools are changing not only what hospitals can build, but who does the building. When a single clinician, program manager, or developer can “10 or 100x their capabilities,” it changes the build-versus-buy equation. He pointed to examples including work with OpenAI on rare disease diagnosis and tools that synthesize ICU information from disparate sources.

Yet Brownstein drew a clear line between building a prototype and deploying a system across an enterprise. “It’s very easy to vibe code a prototype,” he said. “It’s very different to have it deployed in the enterprise.”

That difference is the deployment gap: the distance between a tool that works in a demo and one that is secure, governed, workflow-integrated, and ready to be tested safely.

Operational AI and Clinical Risk Are Not The Same

Bhandari urged leaders to distinguish between operational AI and clinical AI. The upside is large in both, but the risks are not the same: operational tools may automate administrative processes, while clinical tools may directly affect care pathways or patient outcomes. He offered informed consent forms in clinical trials as an example—teams may need to generate many similar but jurisdiction-specific documents, and automating a first draft for human review cuts tedious work while preserving oversight. But when AI touches patient care, he said, leaders need a clear risk framework: whether the tool is in the care pathway, whether it could affect care, and what safeguards are required before deployment. Governance, validation, ethics, bias review, and auditability, he said, are the “veggies” organizations have to eat if they want the “AI cake.”

Natural Language Can Hide Assumptions

Mattie focused on a risk that becomes harder to see as AI tools get easier to use. Natural language feels intuitive, she said, but that is part of the problem: when someone specifies a workflow in plain English, they may not realize they are also encoding assumptions about which patients are included, which data fields matter, and which outcomes are optimized.

In her work on algorithmic bias, the greatest risk is often not a system that obviously fails but one that appears to work until its results are stratified by factors such as insurance status, geography, or race. Bias may be embedded not only in code, but in training data, outcome variables, or the population a prototype was built around.

She urged teams to make certain questions routine: “What population does this perform well on? Where was it tested? What happens in subgroups?” Those questions, she said, “have to become reflexes and not afterthoughts.” For Mattie, equity is not a final feature to add once a product is nearly done; it is a design constraint from the beginning.

Governance Must Meet Bottom-Up Innovation for AI in Healthcare

The panelists did not argue for stopping bottom-up experimentation—they argued the opposite. Because clinicians, researchers, and operators are close to the problems, they should have pathways to test ideas. But those pathways need guardrails.

Brownstein described Boston Children’s AI governance structure, which pairs clinicians with representatives from IT, legal, finance, and compliance to catalog AI efforts, assess risk, prioritize projects, and reduce duplication—supporting innovation while staying careful about which tools move beyond prototype into deployment. Those control functions, the panelists noted, need to be involved early, especially when projects touch personal data or enterprise-scale implementation.

The worst case, Brownstein warned, is a clinician getting far along in building a tool only to have to ask for forgiveness or favors to deploy it. Bhandari framed the same need through experimentation, arguing that AI projects should follow the scientific method: start with a hypothesis, define a plan, measure results, set an endpoint, and iterate. That approach helps organizations build “the muscle of experimentation with guardrails.”

Leaders need build literacy

A recurring theme was that leaders need a new kind of literacy. Mattie called it “build literacy”: not that every clinician or executive needs to code, but that they need enough understanding to be responsible stewards of systems they may help build, evaluate, approve, or deploy. She described three components—problem specification, evaluation, and accountability. Leaders need to define what problem they are solving, for whom, and what success looks like; ask about accuracy, safety, and fairness before piloting; and know who owns the system, who monitors it, and what happens when it fails.

“The goal isn’t for us to produce clinician coders,” Mattie said. “It’s to produce leaders who won’t be fooled by a demo.”

That idea is central to Harvard Chan School’s ALA program, Beyond Vibe Coding: Building AI Solutions to Transform Health Care. The live, online program is designed for leaders ready to move beyond AI strategy and begin building real solutions—safe, bounded, auditable, and evidence-generating ones. Participants bring an opportunity area and work through a structured process: defining the problem, validating the need, scoping the solution, and developing a prototype, while learning how to move from early experimentation to production-ready solutions.

For organizations, the implication is clear: the solution cannot rest on individual development alone. 

Harvard T.H. Chan School of Public Health offersEmerging Women Executives in Health Care, which strengthens the career paths for women in health care.

About The Author


Last Updated

Get the latest public health news

Stay connected with Harvard Chan School