Archive/A Structural Labeling-File Audit and Secondary Model-Output Evaluation of Korean Specialized and Essential Medical Knowledge Datasets for Medical AI Research
A Structural Labeling-File Audit and Secondary Model-Output Evaluation of Korean Specialized and Essential Medical Knowledge Datasets for Medical AI Research
Mi-ae Yang, Kang-Su Ha
22 juillet 2026
en

Abstract

Background: Korean-language medical question-answering datasets are increasingly used for large language model (LLM) development, but structural completeness alone does not establish model performance or clinical validity. We examined two national Korean medical knowledge datasets by combining a public labeling-file audit with a secondary evaluation of item-level outputs distributed with the corresponding official LLM packages. Methods: We reviewed the official documentation and complete public download inventories, parsed 34 training/validation labeling archives containing 31,036 records, and audited required fields, duplication, text length, question type, and specialty distribution. We also audited the official Qwen2.5-14B LoRA model packages and independently recalculated performance from their distributed item-level KorMedMCQA output files (2494 identical items per model). Accuracy was reported with Wilson 95% confidence intervals; model outputs were compared using an exact McNemar test and a paired bootstrap confidence interval. Because no compatible local GPU was available, model inference was not independently rerun. Results: The documentation described 34,487 labeled QA pairs, of which 31,036 (90.0%) were present in the publicly accessible training/validation labeling files; no public test-labeling archive was listed. Required fields were complete, no duplicated qa_id values were found, and one duplicated question-answer pair occurred in the Essential dataset. Multiple-choice items comprised 78.7% of all documented QA pairs, and the top three domains comprised 56.8%. In the distributed KorMedMCQA outputs, the Essential-care model answered 1603/2494 items correctly (64.27%; Wilson 95% CI 62.37–66.13%), while the Specialized-medicine model answered 1596/2494 correctly (63.99%; 95% CI 62.09–65.85%). The paired difference was 0.28 percentage points (bootstrap 95% CI −0.44 to 1.00), with no significant difference by exact McNemar test (46 vs. 39 discordant correct items; p = 0.515). Conclusions: The public training/validation labeling files showed favorable basic structural completeness, but the unavailable test labels, multiple-choice predominance, domain imbalance, limited record-level provenance, and incomplete model-package traceability constrain claims of clinical readiness. The distributed model outputs demonstrated moderate examination-style benchmark performance without a significant difference between the two models. These findings support use as research infrastructure, not evidence of clinical validity, and indicate the need for independent inference reproduction, clinician-led open-ended and safety evaluation, temporal updating, specialty-stratified reporting, and human oversight before clinical use.

IPC Classification

G06A61

Keywords

structurallabeling-fileauditsecondarymodel-outputevaluationkoreanspecializedessentialmedicalknowledgedatasetsresearchbiomedinformaticsbackgroundkorean-languagequestion-answeringincreasinglyusedlargelanguagemodeldevelopmentcompleteness
Citer cette publication

€ 4.00