AI Outperforms Doctors in Patient Care, Tests, and Treatment

Artificial intelligence (AI) is making significant strides in the medical field, with new systems capable of performing a wide range of clinical tasks. These AI systems can ask patients about their symptoms, determine which tests are needed, interpret results, prescribe medication, and even plan hospitalizations. Another AI has been developed to adjust treatment plans based on disease progression and patient responses after multiple visits. These advancements have shown results that are comparable to or even better than those of doctors in certain areas during virtual clinical trials.

The research on two groundbreaking AI systems—“MIRA,” developed by Dresden University of Technology and Heidelberg University Hospital in Germany, and “AMIE,” created by Google Research and Google DeepMind—was published jointly in the international journal Nature on the 17th (local time). Unlike existing medical AIs that typically perform limited tasks such as disease diagnosis or medical imaging analysis, these two systems are designed to support the entire patient management process—from taking medical history to selecting tests, making treatment decisions, and managing follow-up care.

MIRA: AI Capable of Tests, Prescriptions, and Hospitalization Planning

MIRA is an AI agent that operates in an isolated virtual environment modeled after hospital electronic medical records. It interacts with another AI acting as a patient to confirm symptoms, medical history, current medications, allergies, and more. Based on this information, it orders blood, urine, microbial, and imaging tests and makes diagnoses from the results.

MIRA prescribes specific drug names and dosages, selects surgeries or procedures, and decides whether to hospitalize patients. Unlike conventional medical AIs that provide vague answers like “possible pneumonia,” MIRA generates standardized test and prescription orders compatible with electronic medical records. To achieve this, it mimics the operation of hospital electronic medical records by applying international medical coding systems that classify tests, drugs, and diseases.

Researchers evaluated MIRA using 574 patient records from U.S. intensive care and emergency databases, covering eight conditions: appendicitis, cholecystitis, diverticulitis, pancreatitis, pulmonary embolism, urinary tract infection, pneumonia, and pancreatic cancer. The overall diagnostic accuracy was 88.9%. For appendicitis, MIRA correctly identified 146 out of 148 cases (98.6% accuracy), 92.3% for pancreatitis, and 90% for pulmonary embolism. However, performance varied by condition, with 72.4% for pneumonia and 77.6% for urinary tract infections.

MIRA’s average diagnostic accuracy (87.8%) was compared against four specialists evaluating the same 311 cases under identical conditions, outperforming the specialists’ 78.1%. MIRA followed a clinical workflow similar to real emergency rooms, progressing from patient interviews to physical exams, blood tests, imaging, and finally selecting medications or surgeries.

AMIE: AI Adjusting Treatment Plans Based on Patient Responses

Google’s AMIE is designed to track disease changes and treatment responses over multiple outpatient visits. Even after diagnosing a condition, doctors must decide on additional tests, whether to continue or change medications, and when to schedule follow-ups. This dynamic decision-making process, called “management reasoning,” is AMIE’s focus.

AMIE operates with two specialized AIs: one for patient interactions and another for overall treatment planning. Built on Google’s generative AI “Gemini,” it analyzes patients’ prior medical records and current conditions, referencing guidelines from institutions like the U.K.’s National Institute for Health and Care Excellence to inform test and treatment plans.

Researchers created 100 virtual patient cases across cardiology, respiratory, obstetrics/gynecology, gastroenterology, and neurology/musculoskeletal fields. AMIE’s performance was compared against 21 primary care physicians during three simulated visits per patient. Trained actors played patients, and specialists evaluated the dialogue records and treatment plans.

AMIE received appropriateness ratings of 95% for the first visit, 96% for the second, and 98% for the third, compared to physicians’ 72%, 80%, and 81%, respectively. AMIE also scored higher in specificity, detailing not just “prescribe antibiotics” but including drug names, dosages, administration methods, treatment durations, and follow-up timelines.

However, AMIE’s design to generate lengthy, specific answers may have advantaged it in evaluations, and specificity does not guarantee real-world efficacy or safety.

Virtual Trial Success… Further Validation Needed for Real-World Use

Experts caution against interpreting these results as “AI defeating doctors.” MIRA’s evaluation used historical electronic medical records in a simulated environment, with “patients” being AIs programmed to answer based on recorded histories—not real patients who might omit symptoms, provide inconsistent accounts, or experience rapid condition changes.

AMIE’s trial also involved trained actors in text-based simulations, not actual patients. Real-world care involves variables like physical exams, costs, test availability, patient preferences, and interdisciplinary consultations.

Both teams emphasized that AIs are not yet ready for independent patient care. Clinical trials with real patients, safety assessments, and oversight systems to prevent erroneous tests or prescriptions are needed. Responsibility frameworks for harm caused by AI decisions—whether falling on doctors, hospitals, or developers—also require resolution.

While MIRA is an “action-oriented AI” executing tests and treatments within electronic medical records, AMIE is a “disease management AI” adapting plans across multiple visits. These studies highlight how medical AIs are evolving beyond diagnostic tools to broadly support doctors’ clinical workflows.

월 5900원 멤버십, 신문 독자에게는 2900원, 조선멤버십
55000원 상당의 신문-잡지 8종 마음껏 보기, 조선멤버십
현금처럼 쓸 7000포인트 받아 알뜰한 쇼핑, 조선멤버십