의료 인공 일반 지능(MAGI)은 하나의 기반 모델로 다양한 의료 작업을 해결할 수 있으며, 이는 의료 영역에서 매우 실용적입니다. 서로 다른 작업 간에 의학 지식을 충분히 공유함으로써 많은 양의 작업별 데이터 요구 사항을 크게 줄일 수 있습니다. 그러나 제한적이고 복잡한 의료 데이터로 강력하게 일반화할 수 있는 모델을 설계하는 문제로 인해 대부분의 기존 접근 방식은 작업별 모델을 개발하는 경향이 있습니다. MAGI를 향한 한 걸음을 내딛기 위해 MOTOR(Medical-knOwledge-enhanced mulTimOdal pretRaining)라는 새로운 패러다임을 제안합니다. MOTOR에서는 두 종류의 기본 의학 지식, 즉 일반 지식과 특정 지식을 보완적으로 결합하여 일반 사전 교육 과정을 향상시킵니다. 결과적으로, 포괄적인 기본 지식을 갖춘 기초 모델은 더 나은 교차 모달 정렬을 위해 사전 훈련된 방사선 사진 데이터에서 압축 표현을 학습할 수 있습니다. MOTOR는 AI 시스템의 두 가지 핵심 지능인 이해와 생성을 하나의 의료 기반 모델로 통합하여 보다 다양한 의료 업무를 유연하게 처리합니다. 포괄적인 평가를 가능하게 하고 추가 연구를 용이하게 하기 위해 흉부 X-레이 보고서 생성 및 의료 시각적 질문 답변과 같은 광범위한 다운스트림 작업을 포함하는 의료 다중 모달 벤치마크를 구성합니다. 우리의 벤치마크에 대한 광범위한 실험은 MOTOR가 간단한 작업 지향 적응을 통해 유망한 결과를 얻는다는 것을 보여줍니다. 시각화는 주입된 지식이 의료 데이터의 주요 정보를 성공적으로 강조 표시함을 보여줍니다. MOTOR의 뛰어난 해석성을 보여줍니다. 당사의 MOTOR는 "의대생"이 되는 인간의 실습을 성공적으로 모방하여 "전문가"가 되는 과정을 가속화합니다.


https://arxiv.org/abs/2304.14204

Towards Medical Artificial General Intelligence via Knowledge-Enhanced Multimodal PretrainingMedical artificial general intelligence (MAGI) enables one foundation model to solve different medical tasks, which is very practical in the medical domain. It can significantly reduce the requirement of large amounts of task-specific data by sufficiently sharing medical knowledge among different tasks. However, due to the challenges of designing strongly generalizable models with limited and complex medical data, most existing approaches tend to develop task-specific models. To take a step towards MAGI, we propose a new paradigm called Medical-knOwledge-enhanced mulTimOdal pretRaining (MOTOR). In MOTOR, we combine two kinds of basic medical knowledge, i.e., general and specific knowledge, in a complementary manner to boost the general pretraining process. As a result, the foundation model with comprehensive basic knowledge can learn compact representations from pretraining radiographic data for better cross-modal alignment. MOTOR unifies the understanding and generation, which are two kinds of core intelligence of an AI system, into a single medical foundation model, to flexibly handle more diverse medical tasks. To enable a comprehensive evaluation and facilitate further research, we construct a medical multimodal benchmark including a wide range of downstream tasks, such as chest x-ray report generation and medical visual question answering. Extensive experiments on our benchmark show that MOTOR obtains promising results through simple task-oriented adaptation. The visualization shows that the injected knowledge successfully highlights key information in the medical data, demonstrating the excellent interpretability of MOTOR. Our MOTOR successfully mimics the human practice of fulfilling a "medical student" to accelerate the process of becoming a "specialist". We believe that our work makes a significant stride in realizing MAGI.arxiv.org