Realizing the Promise: Artificial Intelligence in Endocrinology

Artificial intelligence (AI) is steadily moving from novelty to necessity in medicine, and endocrinology is no exception. The ENDO 2026 session “Artificial Intelligence in Endocrinology: Practical Uses, Lessons Learned, and What Comes Next” showed attendees that AI is here, and those who wait too long to engage with it may find themselves playing catch-up.

At a packed session on Saturday, June 13th at ENDO 2026, two researchers who study artificial intelligence (AI) in healthcare set out to answer some of the many nagging questions related to the purported AI revolution in “Artificial Intelligence in Endocrinology: Practical Uses, Lessons Learned, and What Comes Next.” Says session chair Juan P. Brito, MD, MSc, professor of medicine and director of the Care and AI Laboratory at Mayo Clinic, as well as medical director of the Shared Decision Making National Resource Center and innovation and quality chair for Mayo’s Division of Endocrinology: “It was an opportunity to connect two complementary perspectives, one focused on practical AI tools that clinicians can use today and the other on the lessons learned from implementing AI in real healthcare settings.”

David Toro-Tobón, MD, a Mayo Clinic Scholar and assistant professor of medicine in the Division of Endocrinology, Diabetes, Metabolism and Nutrition at Mayo Clinic in Rochester, Minn., spoke first, offering a very practical presentation that showed how AI can improve day-to-day clinical and academic work. “Rather than focusing on futuristic applications,” says Brito, “he demonstrated tools that many endocrinologists can begin using immediately to improve efficiency while recognizing their limitations.”

Naykky M. Singh Ospina, MD, professor in the Division of Endocrinology, Diabetes, and Metabolism at the University of Florida in Gainesville complemented that perspective by discussing the final stages of the AI development and implementation pathway. “She presented two illustrative examples: The implementation of AI for diabetic retinopathy screening, where AI has already demonstrated clinical value in practice, and the use of AI for thyroid ultrasound assessment, highlighting both its potential and its current limitations,” explains Brito.

“Most of us have been hearing how AI can be really disruptive in medical practice for a few years now,” says Singh Ospina. “The session was an exercise in reflecting on what has actually changed, what might be changing right now, and what is needed for that promise of AI to actually improve the care of our patients in clinic. We wanted to answer, what can you use AI for right now? Is this the disruption we were expecting?”

Brito says the invitation to chair tracked closely with his own research, which sits at “the intersection of endocrinology, evidence-based medicine, and AI,” with a focus on “evaluating how AI can improve clinical decision-making and patient care while remaining centered on evidence and implementation.” Camilo Daniel Gonzalez, MD, PhD of the Autonomous University of Nuevo León in Monterrey, Mexico, also served as chair of the session.

Turning a Would-be Threat into an Opportunity

Toro-Tobón’s portion focused on appropriate use of what he refers to as “Everyday AI” — the different types of chatbots and generative AI clinicians use at different stages of work, like AI scribes; decision support at the point of care, like OpenEvidence and other tools for clinical decision aid or data retrieval; and administrative support for documentation or for research and academic work. “There’s this idea that these tools are coming for us, coming for our place, but I don’t think that’s true,” he says. “Fundamentally, artificial intelligence works differently than human intelligence. They have pitfalls and potentials that are complementary. So, I see these tools as something that works together with us, not in replacement of us.” Used well, he argues, AI hands back something clinicians are chronically short on: time. “If you’re using it the right way and offloading the cognitive load of tasks that don’t really need you and can be done by a machine, you can redirect all that time you’re saving to actual patient care: looking at your patients, understanding them, doing the things you’re passionate about.”

Of course, this also cuts the other way: “If we don’t understand the risks of using AI, and the way it should be appropriately introduced into the clinic, then we risk making things worse: a higher workload from not using it appropriately, or potential patient harm if we’re not careful about hallucinations and other things that come with it. The stakes are much higher, and the level of error we can accept can’t be as lenient. That’s what makes it challenging.” Some of the key principles to address this challenge include:

  • Passive interaction with generative AI risks cognitive deskilling by creating an illusion of expertise before genuine clinical judgment is earned.
  • Trainees must dynamically match their AI interaction style to the clinical risk level, using a “Centaur” approach with heavy human oversight for high-stakes decisions.
  • Core clinical reasoning must be maintained as an autonomous, tech-free competence to ensure clinicians possess the foundational expertise required to effectively audit AI outputs.

Toro-Tobón’s own research bridges the research and clinical worlds. A substantial portion of his academic work focuses on the intersection of AI and thyroid diseases, and healthcare delivery, and he is completing a master’s degree in AI in healthcare at Mayo Clinic. “I have a little bit of dual expertise: developing tools that help researchers better understand and treat thyroid disease as well as healthcare delivery and AI — how using these tools play out in the clinic,” he says. Indeed, part of what motivated his pursuit of a master’s in AI was the widening gap between people who understand the data science and people who understand the clinic, with too little communication between them. “There are a lot of people with very extensive data science knowledge who don’t have clinical expertise, building AI for healthcare,” he says. “And there are a lot of people with phenomenal clinical expertise who, because we’re in this AI boom, are trying to build AI applications themselves.” The result can be tools that do not fit their purpose as well as complacency about perceived outcomes. “There’s a lot of enthusiasm, and a lot of papers being submitted to our journals, but when you look closely, there’s a problem with AI spoofing: studies demonstrating phenomenal performance but with significant methodological flaws.” The papers themselves may be legitimately published, he is careful to note, but the caveats may be lost on the individual reader. “What they’re reading isn’t practice-changing today,” says Toro-Tobón.

“AI is not going to replace us. Instead … clinicians who work with AI are going to replace those who don’t. These tools are here to augment us, as long as we’re using them the right way, at the right time. They have the potential to help us do our job better, give us time back to interact with our patients, and do what we feel passionate about. So, the most important thing is not to feel threatened by AI, but to embrace it in a way that’s safe, appropriate, and ethical and that always keeps the patient at the center.” – David Toro-Tobón, MD, Mayo Clinic Scholar, assistant professor of medicine, Division of Endocrinology, Diabetes, Metabolism, and Nutrition, Mayo Clinic, Rochester, Minn.

The goal of the Mayo Clinic master’s program is to graduate clinician-scientists with technical fluency on both sides. Toro-Tobón explains that, “it treats AI as a basic science, a new technique we have at our disposal for clinical questions we already have, a new tool for our research.” The curriculum also covers implementation and ethics, because the cost of getting it wrong in medicine could be catastrophic. Take data privacy, one of the most common threads of the discussion. “There were certainly concerns about privacy, and how to use these tools in a way that’s compliant with regulations,” says Toro-Tobón. A core distinction he tried to drive home is that a tool that is safe for a clinician’s personal use or for general professional use with business-confidential information is not automatically safe for patient data. “Using your personal data or business-confidential information on an AI tool is not the same as using patient data, which is HIPAA-regulated. The tools you use with patient information have to be specifically reviewed and approved for that purpose, and that adds another layer of complexity.”

To illustrate just how easy that line is to miss, he offered a real-world example. Some AI tools are embedded in the software many institutions already use, and logging in with a work account produces a green shield with reassuring language about business data protection. “There’s a big misconception there,” warns Toro-Tobón, “and many health systems have even told their providers, ‘you can use this; it’s HIPAA-compliant.’ But that’s not true. It’s business-compliant, meaning they’re not using that data to train the models, but which doesn’t mean it meets all the criteria to be HIPAA-compliant.”

Whether such a tool can be used with patient data depends on the specific product configuration, institutional approval, contractual terms, and whether the appropriate privacy and security arrangements are in place. “This is one of those details that, if you’re not aware of it, you might be putting your patients’ information at risk without realizing it.” With other available tools that are HIPAA-compliant by design (e.g., Doximity Ask and OpenEvidence), users sign in with a National Provider Identifier number that establishes an agreement to handle patient data in a HIPAA-compliant way, but that protection is personal, not institutional, and is why some health systems opt out of it, explains Toro-Tobón. That data does not technically belong to the individual; it belongs to the institution. “We’re in healthcare, which is a different beast,” he says. “We need to be very safe when we use these tools. Good isn’t good enough; these tools have to be exceptional for us to use them in healthcare.”

Gaps Preventing Widespread AI Adoption

These lessons are particularly relevant to Singh Ospina’s research, which centers on patient-centered care and implementation science, with a focus on the patient, clinician, and healthcare system factors that shape the implementation of innovations, including AI solutions, in endocrinology. If Toro-Tobón’s section explored what we can and cannot use today and how, Singh Ospina’s was to reflect on why more of the rest of the AI promise has yet to materialize.

She organized this gap into three pillars. The first is technical: “We still need the AI technology to mature for the kinds of questions we ask in medicine, which are often dependent on context, often longitudinal, and often involve a lot of uncertainty,” she says. The second is clinical: “Even if a tool can make a prediction or suggest the risk of a particular outcome, we need to make sure it’s actually improving patient outcomes — that’s what we need the tools to show or do.” The third is implementation: getting patients, clinicians, and health systems to actually adopt and trust a tool once it exists. The three pillars obviously create a circular problem, in which progress on the third pillar is gated by progress on the first two, and progress on the first two is hard to justify without traction on the third. “You need evidence of clinical outcomes and trust in the technology to be able to implement it. So, AI becomes both a technical and a social challenge,” explains Singh Ospina.

To illustrate her points, she did a head-to-head comparison of an AI tool that has cleared all three hurdles and one that has not, at least not yet. The success story is a U.S. Food and Drug Administration (FDA)-approved imaging tool for diabetic retinopathy that mitigates certain patient barriers to the need for regular retinal screening, such as transportation to multiple appointments and scheduling constraints. “The AI tool takes a picture in clinic, the system flags high-risk findings anything suspicious on the spot, and the patient ideally leaves with a follow-up already arranged,” says Singh Ospina. She adds that the tool alone is not what moves the needle. “You need the tool plus integration into the workflow, making sure the patient actually gets the appointment.”

Naykky Singh Ospina, MD, Assistant Professer, DOM-Endocrinology

“We still need the AI technology to mature for the kinds of questions we ask in medicine, which are often dependent on context, often longitudinal, and often involve a lot of uncertainty. Even if a tool can make a prediction or suggest the risk of a particular outcome, we need to make sure it’s actually improving patient outcomes — that’s what we need the tools to show or do.” – Naykky M. Singh Ospina, MD, professor, Division of Endocrinology, Diabetes, and Metabolism, University of Florida, Gainesville, Fla.

The retinopathy tool checks every box: it works (technical); it improves adherence to follow-up (clinical outcome); and it is backed by FDA approval, clinical-guideline endorsement, and even specific billing codes (implementation). The contrast case is thyroid ultrasound. Several AI tools now exist that can analyze a thyroid ultrasound and flag how suspicious a nodule looks for cancer, but they have not demonstrated improvements in care. “Some clinicians might benefit from using these tools and others might not,” she says. With the clinical gap unresolved, the implementation pieces will not click into place, leaving us without guideline endorsement or quality metrics, for example. Moreover, many of the models are trained predominantly on the most common type of thyroid cancer, raising clinician concern about their accuracy for patients with rarer subtypes.

Part of the difference in what succeeds comes down to the nature of the question each tool is being asked to answer, explains Singh Ospina. Retinopathy screening is binary (problems, yes or no for high-risk actionable findings?), whereas thyroid ultrasound comes with more uncertainty. “There’s a need to contextualize the ultrasound findings for individual patients and decide what to do next,” she says. “The technology we have so far can be very helpful with narrow-scope questions. But what we want, and where the real revolution will occur, is technology that can help us with more complex decisions.”

At least part of the solution to this problem, according to Singh Ospina, is to involve all stakeholders from the outset. “You need multidisciplinary teams with experts across all three pillars, from inception,” she says. The diabetic retinopathy tool succeeded in part because it was designed around an actual point of friction in care — the wait, the delay, the burden of getting screened and getting an appointment — rather than around what the technology was capable of in the abstract. Thyroid ultrasound AI faces a steeper climb for the opposite reason: The underlying clinical problem is more complex, and adding an AI agent into that decision raises new questions about when and how it should weigh in. “It’s not a reason to abandon the effort,” she says, “but a reason to keep clinicians and patients embedded in the design process from day 1 rather than brought in once the tool is already built.” Researchers are working to close that distance by incorporating patient demographics, values, and preferences, but more must be done.

While David Toro Tobon, MD (far left) is speaking at the podium, co-chairs (l to r)  Juan P. Brito, MD, and Camilo D. Conzalez, MD, and fellow speaker Naykky Maruquel Singh Ospina, MD, listen.

The Takeaways

Both speakers described an engaged audience, and Brito came away with a similar read. “Their questions reflected an increasingly sophisticated understanding of AI, moving beyond curiosity about the technology itself toward thoughtful discussions about evaluation, implementation, and patient impact,” he says. “That, to me, reflects the growing maturity of our field’s approach to AI.”

For Toro-Tobón, the energy in the room is part of what keeps him coming back to ENDO each year. “What I enjoy most about ENDO, this year included, is the level of curiosity,” he says. “That interaction with people who are curious and passionate about what they do sparks discussion, and sometimes you get as much out of that as from the sessions themselves.” He most wanted his particular audience to walk away with continued curiosity about where and how AI can help us. “AI is not going to replace us. Instead, what’s going to happen is that clinicians who work with AI are going to replace those who don’t,” he says. “These tools are here to augment us, as long as we’re using them the right way, at the right time. They have the potential to help us do our job better, give us time back to interact with our patients, and do what we feel passionate about. So, the most important thing is not to feel threatened by AI, but to embrace it in a way that’s safe, appropriate, and ethical and that always keeps the patient at the center.”

“While AI holds tremendous promise to improve healthcare, translating that promise into routine clinical practice remains challenging. Methodological issues, implementation barriers, governance, and the need for rigorous evaluation still need to be addressed before many AI applications can be widely and safely adopted.” – Juan P. Brito, MD, MSc, professor of medicine, director, Care and AI Laboratory; medical director, Shared Decision Making National Resource Center; innovation and quality chair, Division of Endocrinology, Mayo Clinic, Rochester, Minn.

Singh Ospina’s closing case was an exhortation: Get clinicians, AI developers, and implementation scientists all involved early, and often, in how these tools get built in a collaborative way. “You can have the best person developing the best AI architecture and knowing exactly what data is needed, alongside clinicians, patients, and health system representatives who understand the reality people are facing,” she says, “and you codesign with implementation in mind, rather than building the tool first and asking afterward who’s going to use it and how.” For now, the diabetic retinopathy success story remains the field’s best proof of concept that the approach works. The rest of endocrinology’s AI promise is still waiting on its turn.

The throughline of both talks, as Brito summarizes it, is that “while AI holds tremendous promise to improve healthcare, translating that promise into routine clinical practice remains challenging. Methodological issues, implementation barriers, governance, and the need for rigorous evaluation still need to be addressed before many AI applications can be widely and safely adopted.” Brito’s own takeaway lands close by. “AI remains a work in progress,” he says. “There is tremendous excitement about its potential, but there is also a growing appreciation for its current limitations and the work that remains before many applications can be routinely integrated into clinical care.” In the near term, he expects that value to keep showing up in the unglamorous parts of the job — “routine cognitive and administrative tasks that improve efficiency” — while AI tools that directly influence diagnosis and treatment decisions “will require additional research, careful evaluation, and thoughtful implementation before they become part of everyday endocrine practice.” That distinction showed up across ENDO 2026 well beyond this one session: “It was exciting to see the field moving from exploration toward thoughtful implementation,” while also “emphasizing the importance of maintaining rigorous standards for evaluating these technologies before they become part of routine practice.”

Taken together, Brito says the two presentations “provided a balanced view of both the immediate opportunities and the longer-term challenges of AI in endocrinology.”

  • Horvath is a freelance writer based in Baltimore, Md. She wrote “The Family Business,” about the father and son pair of endocrinologists, Dan and Phillip Dumesic, in the July issue.


Share this article