Reporting checklist for foundation and large language models in medical research (REFINE): an international consensus guideline | 경쟁 | 위협 3/3 | 2026 | Diagn Interv Radiol | 개발군 57명/17개국, 패널 54명/16개국이 R1·R2 완료. 최종 44항목 6섹션 | 게재지가 영상의학 전문지다. 일반 의료 LLM 커뮤니티에서의 채택률은 미지수이며, 채택 실태 자체가 연구 대상이 될 수 있다. |
Structured taxonomy and framework for developing medical benchmark in large language models derived from scoping review (READY) | 경쟁 | 위협 3/3 | 2026 | NPJ Digit Med | 벤치마크 55편. 도메인 전문가 5명이 독립 적용해 일관된 평가자 간 일치도 확인 | 대상이 벤치마크 데이터셋이다 — 연구도, 근거도 아니다. 그리고 축이 정적이다: model versioning·drift, API 변경, stochasticity, prompt sensitivity, contamination, benchmark saturation, LLM-as-judge 편향, agent/RAG/tool-use, 비용·지연, 종단 평가, closed model 재현성이 중심축에 없다. 저자들 스스로 Discussion에서 "LLM으로 벤치마크를 만드는 방법론 자체가 아직 미성숙하고, 기존 벤치마크 연구 간 표준 보고 관행이 부족하다"고 인정한다. |
Navigating the landscape of medical artificial intelligence reporting guidelines | 경쟁 | 위협 3/3 | 2025 | Lancet Digit Health | 해당 없음 | 지형을 서술하는 것과, 그 지형에서 무엇이 빠졌는지를 데이터로 보이는 것은 다르다. 후자는 아직 열려 있을 가능성이 있다. |
The TRIPOD-LLM reporting guideline for studies using large language models | 경쟁 | 위협 3/3 | 2025 | Nat Med | 19개 main item + 50개 subitem. 전 범주 공통 14 main / 32 subitem | living document라고 스스로 선언했다는 것은, 갱신의 근거를 제공하는 연구에는 자리가 있다는 뜻이다. 즉 대체가 아니라 입력이 되는 경로. |
Analyzing evaluation methods for large language models in the medical field: a scoping review | 방법틀 | 위협 3/3 | 2024 | BMC Med Inform Decis Mak | 142편 | 검색 기간이 2023년 9개월뿐이고 DB 3개다. GPT-4 이후, agent·RAG·멀티모달 시대는 전혀 포함되지 않았다. 2티어 게재. |
Holistic evaluation of large language models for medical tasks with MedHELM | 경쟁 | 위협 2/3 | 2026 | Nat Med | 5개 범주 / 22개 하위범주 / 121개 태스크. 37개 평가. 프런티어 LLM 9종 비교 | 태스크 taxonomy이지 보고 항목 taxonomy가 아니다. 무엇을 평가할지를 정리했지, 평가를 어떻게 보고할지는 다루지 않는다. |
Reporting Guideline for Chatbot Health Advice Studies: The CHART Statement | 경쟁 | 위협 2/3 | 2025 | JAMA Netw Open | Delphi 이해관계자 531명, 동기 패널 48명. 최종 12항목 39subitem | 범위가 챗봇 건강조언에 한정된다. EHR 연동·agent·retrieval·멀티모달은 범위 밖. |
Reporting guideline for the use of Generative Artificial intelligence tools in MEdical Research: the GAMER statement | 경쟁 | 위협 2/3 | 2025 | BMJ Evid Based Med | 26개국 전문가 51명 참여(Delphi 설문 44명). 최종 9항목 | 연구 도구로서의 GAI 사용 보고이지, 연구 대상으로서의 의료 LLM 평가 보고가 아니다. |
Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review | 분모 | 위협 2/3 | 2025 | JAMA | 519편 | 2024년 2월까지다. 그 이후의 agent·retrieval·멀티모달 평가와 LLM-as-judge 확산은 잡히지 않았다. 또 "무엇이 부족한지"는 밝혔지만 보고 항목 수준의 처방은 내지 않았다. |
Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks | 분모 | 위협 2/3 | 2025 | J Med Internet Res | 3,917건 스크리닝 → 39개 벤치마크. 총 230만+ 문항, 45개 언어, 172개 전문과 | 벤치마크를 단위로 했기 때문에, 개별 평가 연구의 절차 보고(반복측정·채점자·프롬프트)는 대상이 아니다. |
A framework for human evaluation of large language models in healthcare derived from literature review (QUEST) | 선례 | 위협 2/3 | 2025 | NPJ Digit Med | 142편 리뷰 | 인간 평가에 한정된다. LLM-as-judge, 자동 평가, agent 평가, 재현성·버전 관리 같은 기술 요소는 대상이 아니다. |
Large Language Models for Chatbot Health Advice Studies: A Systematic Review | 선례 | 위협 2/3 | 2025 | JAMA Netw Open | — | 챗봇 건강조언 한정. |
Automating expert-level medical reasoning evaluation of large language models (MedThink-Bench) | 방법틀 | 위협 1/3 | 2026 | NPJ Digit Med | 500개 고난도 문항, 10개 의학 도메인, 전문가 작성 단계별 rationale 동반 | 단일 벤치마크. 이런 상관계수·시간 절감 보고가 표준화되어 있지 않다. |
What to Report? A Systematic Review of Medical AI Reporting Guidelines: Preliminary Results | 경쟁 | 위협 1/3 | 2025 | Stud Health Technol Inform | 9편 — 임상·비임상 DB 양쪽 | 9편 예비 결과. 전수 규모, 기술 요소 축의 부재를 정량화하지 않았다. 프로시딩이라 1티어 인용 무게가 가볍다. |
FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare | 토대 | 위협 1/3 | 2025 | BMJ | — | 원칙 수준이라 LLM 평가의 구체적 방법 항목(반복측정, 채점자 수, 프롬프트 기록)까지 내려가지 않는다. |
A systematic review of large language model (LLM) evaluations in clinical medicine | 분모 | 위협 1/3 | 2025 | BMC Med Inform Decis Mak | 761편 | 분류 수준에 머무르고 보고 항목 처방으로 가지 않는다. |
Standardizing and Scaffolding Health Care AI-Chatbot Evaluation: Systematic Review (HAICEF) | 방법틀 | 위협 1/3 | 2025 | JMIR AI | 266건 → 152건 스크리닝 → 전문 21건 → 프레임워크 11개 포함. 질문 356개 → 중복제거·관련성 검토 후 271개 | 챗봇 한정, JMIR AI 게재. |
Evaluating clinical AI summaries with large language models as judges | 방법틀 | 위협 1/3 | 2025 | NPJ Digit Med | 실제 EHR 다문서 요약 | 요약 태스크 하나. 그리고 이런 검증 절차(ICC·CI·인간 대조)를 보고하도록 요구하는 지침은 아직 없다. |
Update to the PRISMA guidelines for network meta-analyses and scoping reviews and development of guidelines for rapid reviews: a scoping review protocol | 진입구 | 위협 1/3 | 2025 | JBI Evid Synth | PRISMA-NMA·PRISMA-ScR 개정 + PRISMA-RR 신규 개발. 언어 제한 없음, 2인 독립 스크리닝·추출 | 프로토콜이다 — 아직 결과가 나오지 않았다. 그리고 포함 기준이 "보고 완전성을 평가한 연구"다. 즉 내가 의료 LLM 리뷰의 보고 완전성을 실측하면 그 자체가 이 개정의 입력 자격을 갖는다. |
Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review | 선례 | 위협 1/3 | 2024 | Lancet Digit Health | AI RCT 86건 | RCT만 대상. LLM 평가 연구 대다수는 RCT가 아니다. |
Artificial intelligence for surgical scene understanding: a systematic review and reporting quality meta-analysis | 선례 | 위협 0/3 | 2026 | NPJ Digit Med | 188편 | 수술 장면 이해라는 좁은 영역. |
The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence | 토대 | 위협 0/3 | 2025 | Nat Med | — | 진단 정확도 설계에 한정. |
A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation | 선례 | 위협 0/3 | 2025 | NPJ Digit Med | — | 요약 한정. |
TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods | 토대 | 위협 0/3 | 2024 | BMJ | — | 생성형·대화형 모델의 평가 절차(프롬프트, 반복측정, 채점자)는 대상이 아니다. |
A critical assessment of using ChatGPT for extracting structured data from clinical notes | 방법틀 | 위협 0/3 | 2024 | NPJ Digit Med | 폐암 병리보고서 1,000+건, 소아 골육종 191건 | 단일 연구의 좋은 관행이지, 요구사항이 아니다. |
Conceptualizing the reporting of living systematic reviews | 방법틀 | 위협 0/3 | 2023 | J Clin Epidemiol | — | 변하는 것이 "근거"이지 "개입 자체"가 아니다. 새 논문이 나오는 것을 다루지, 평가 대상 모델이 사라지거나 교체되는 상황은 다루지 않는다. 바로 그 지점이 의료 LLM의 문제다. |
Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI | 토대 | 위협 0/3 | 2022 | Nat Med | — | CDSS 초기 임상 평가에 한정. |
Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension | 토대 | 위협 0/3 | 2020 | Nat Med | — | RCT 형식에 묶여 있어, 시험이 아닌 벤치마크·오프라인 평가는 범위 밖. |
Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies | 방법틀 | 위협 0/3 | 2020 | BMJ | — | 2019년까지, 영상 딥러닝 한정. |
PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation | 토대 | 위협 0/3 | 2018 | Ann Intern Med | 20개 필수 항목 + 2개 선택 항목. Delphi 1차 31명 참여, 최종 24명이 3차 완료 | 2018년 프레임워크다. 문헌 한 편을 안정적인 최소 분석 단위로 전제한다. 같은 이름의 기술이 같은 개입이라고 암묵적으로 가정하며, 모델 버전·접근 시점·프롬프트·추론 설정 같은 실험 provenance를 요구하지 않는다. |
Guidelines for Developing and Reporting Machine Learning Predictive Models in Biomedical Research: A Multidisciplinary View | 토대 | 위협 0/3 | 2016 | J Med Internet Res | — | LLM 이전 시대. |