Abstract:To compare the accuracy of domestic large language models (LLMs) in simulated clinical anesthesia emergency scenarios and to explore their application in improving clinical anesthesia emergency capabilities. Methods Questions related to clinical anesthesia emergencies were submitted to six mainstream domestic LLMs, including Doubao, ERNIE, Yuanbao, DeepSeek, Kimi and Qwen, and corresponding answers were collected. Highperformance LLMs were selected based on the accuracy of responses. A total of 40 anesthesiologists were enrolled and randomly divided into an experimental group (n = 20) and a control group (n = 20). All participants completed a questionnaire based on anesthesia emergency cases before training, and scores were calculated using unified scoring criteria. The experimental group received LLM-assisted training, while the control group adopted a traditional training mode. The training period lasted for one week, with a daily training duration of 120 ± 10 minutes. After training, another questionnaire involving different anesthesia emergency cases was administered for evaluation and scoring. The scores of the two groups were compared to assess the effectiveness of LLM-assisted training. The chi-square test was used to analyze differences in the performance of LLMs in single-choice questions; the Mann-Whitney U test was applied to compare their performance in open multiple-choice questions and overall outcomes. The independent-samples t test and Mann-Whitney U test were used to evaluate the effect of LLM-assisted training. Bonferroni correction was applied to control Type I errors. The significance level was set at α =0.05, and a corrected P < 0.05 was considered statistically significant.Results Doubao demonstrated the best performance in single-choice questions involving basic knowledge, related professional knowledge and specialized knowledge, with an average accuracy of 92.3%, and ranked first with a score of 77.5 in professional practical knowledge (open multiple-choice questions). The experimental group showed significantly greater score improvement than the control group (P < 0.05). All participants in the experimental group passed the post-training test with higher score stability, while 30% of anesthesiologists in the control group showed negative improvement in academic performance. Conclusion Mainstream domestic LLMs have attained desirable maturity and stability in clinical anesthesia emergency practice. Among them, Doubao has leading advantages in knowledge depth, applicability and accuracy. LLM-assisted training can significantly improve anesthesiologists' anesthesia emergency capabilities, with the strengths of prominent efficacy, stable training effects and low adverse risks. LLMs possess favorable application value in the field of anesthesia emergency management.