More content to come, please stay tuned.

From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLM

Chanyong Luo3,a, Jirui Dai5,a, Zhendong Wang4,a, Jing Wang7, Kuichen10, Jiaxi Yang9, Bingjie Lu11, Jiaxin Hao6, Bing Li6, Ruiyang He6, Yiyu Qiao6, Chenkai Zhang8, Kaiyu Wang8, Zhi Liu2,✉, Zeyu Zheng8,✉, Yan Li6,✉, Xiaohong Gu1,✉
1.The School of Chinese Medicine, the Beijing University of Chinese Medicine, Beijing, China
2.The School of Pharmacy, Nanjing University of Chinese Medicine, Nanjing, China
3.The Infectious disease department, Dongfang Hospital, Beijing University of Chinese Medicine, Beijing, China
4.The Gulou Hospital of Traditional Chinese Medicine of Beijing,Beijing, China
5.The Department of Computer Science, Johns Hopkins University, Baltimore, USA
6.The Department of Education, Dongzhimen Hospital, Beijing University of Chinese Medicine, Beijing, China
7.The Department of Pediatrics, Wangjing Hospital, China Academy of Chinese Medical Sciences, Beijing, China
8.The School of Information Engineering, Huzhou University‌, Huzhou, China
9.Research Center for Scientific Data Hub, Zhejiang Lab, Hangzhou, China
10.The Frontier Basic Research Center, Zhejiang Lab, Hangzhou, China
11.The Research Center for High Efficiency Computing Infrastructure, Zhejiang Lab, Hangzhou, China
Med-Shicheng achieves SOTA in multi-master personalized knowledge and skills TCM, Outperforming Top General LLMs.

Abstract

Medicine is fundamentally an empirical body of knowledge accumulated through long-term observation, validation, and refinement, and ultimately realized in the messy, high-variance real clinical practice. Physicians’ diagnostic-and-therapeutic competence is gradually forged through repeated cycles of “application–reflection–improvement,” crystallizing into distinctive, individualized methodologies. Considering the fact of the substantial variation in treatment outcomes, the formation of master physicians’ knowledge systems is necessary but time-consuming and their transmission is often limited in scope, making high-quality expertise difficult to scale and disseminate, and the scarcity of advanced clinical resources. To mitigate these challenges, we propose Med-Shicheng, a general framework that enables large language model to systematically learn and transfer distinguished physicians’ diagnostic-and-therapeutic philosophy, and case-dependent adaptation rules in a standardized manner. Built upon Tianyi, Med-Shicheng comprises five elaborately designed five stages. We take five National Masters of Chinese Medicine/distinguished TCM physicians as representative targets, curate and organize multi-source materials and train a single model to simultaneously internalize the five distinctive knowledge systems across seven tasks: etiology–pathogenesis analysis, syndrome diagnosis, determination of treatment principles, prescription generation, prescription explanation, symptom evolution with corresponding regimen modifications, and clinical advice. Implemented on Qwen2.5-1.5B-Base, Med-Shicheng can be deployed on resource-constrained GPUs, while demonstrating evaluation performance comparable to DeepSeek-R1 and GPT-5. Besides, we further analyze the reliability of LLM-as-a-Judge against physician-based assessment. It shows that although automated judging captures broad performance trends, it exhibits noticeable biases in fine-grained, individualized clinical distinctions, indicating that physician involvement remains necessary when ground truth is unavailable and that judge models require targeted medical domain adaptation to achieve reliable evaluation.

Med-Schiehng TCM heritage Framework

We propose Med-Shicheng, a general paradigm that enables lightweight LLMs to inherit the diagnostic-and-therapeutic (D&T) expertise of distinguished TCM doctors. The framework contains five progressive stages: (1) building a medical in-domain foundation model, (2) learning regular TCM D&T strategies, (3) reasoning-enhanced D&T fine-tuning, (4) fine-tuning on each doctor’s unique medical knowledge, and (5) reinforcement-style refinement of their distinctive reasoning and treatment strategies. For every stage, we curate task-specific datasets that follow TCM clinical logic and cover the full chain from etiology analysis to prescription and follow-up adjustment. This staged design turns a general-purpose LLM into a doctor-aware medical agent that can internalize multiple masters’ styles within a single model and switch behaviors according to the target physician.For more detailed information, please see our paper.

Med-Schiehng TCM heritage Framework illustration

Evaluatino Benchmark

Evaluation framework illustration

Our evaluation framework combines automated LLM-as-a-judge scoring with expert TCM assessment. For automatic evaluation, we use state-of-the-art general LLMs, DeepSeek-V3.2 and GPT-5, as judges. Each model’s response to a clinical case is scored twice, once by each judge, who compare the generated answer with the gold label and decide how clinically desirable it is. Because every task response is a long, structured text spanning seven components, classical token-level metrics such as accuracy, recall or F1-score cannot reflect quality; holistic judging by strong LLMs is more suitable. To complement this, we conduct human evaluation with 18 senior TCM physicians, grouped around the five target masters they directly studied under. Using Delphi consensus, we provide unified criteria and ask them to rate all models along five dimensions: (1) similarity to the target master’s D&T style, (2) consistency with TCM philosophical principles, (3) safety and risk control, (4) therapeutic completeness, and (5) clinical coherence.

Automatic Evaluations

Automatic evaluation by DeepSeek-V3.2 and GPT-5 shows broadly consistent trends across all target TCM masters. One notable exception is HuatuoGPT2-7B on Dr. Huaitang Du’s cases, where it achieves unusually high scores for prescriptions and treatment principles by reproducing label-like content, suggesting possible training-data overlap. Overall, models naturally split into two tiers. The high-performing group includes Med-Shicheng, GPT-5, DeepSeek-R1, Qwen3-235B-A22B-Thinking, and Gemini-2.5-Pro; the lower-performing group includes Qwen2.5-1.5B-Instruct, Tianyi, and HuatuoGPT2. Med-Shicheng is the only lightweight model in the top tier, delivering content quality comparable to hundred-billion-parameter systems. However, no single model dominates every doctor or metric. Using DeepSeek-V3.2 as judge, DeepSeek and GPT-5 often rank in the top two, but Med-Shicheng, Gemini-2.5-Pro and Qwen3-235B-A22B-Thinking sometimes outperform them, for example on the cases of Fengchun Wang, Shaoqin Zhao and Huaitang Du.

Automatic evaluation results illustration

Human Doctors Evaluations

Human doctor evaluation results illustration

Human evaluation reveals a different picture from LLM-as-a-judge scoring. We ask 18 senior TCM physicians to rate all models on each master’s cases, along five dimensions: similarity to the target physician, adherence to TCM theory, safety, therapeutic completeness, and clinical coherence. Overall, human doctors assign Med-Shicheng higher and more stable scores for most masters, including Huaitang Du, Xiaohong Gu, Fengchun Wang and Bowei Qin, whereas GPT-5 and DeepSeek-V3.2 do not consistently place it near the top. This suggests that very large general LLMs can judge long, complex answers under well-designed prompts, but struggle with fine-grained distinctions in individualized clinical style. They tend to evaluate with generic TCM knowledge, even for widely documented experts such as Dr. Du, while human physicians can recognize subtle signatures of each master’s system. The gap is especially visible for Bowei Qin’s cases, where clinicians rank Med-Shicheng near second, but LLM judges push it down to around fourth.

BibTeX

@article{medshicheng2026,
  title={Med-Shicheng:From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLM},
  author={First Author and Second Author and Third Author},
  year={2026},
  eprint={COMING SOON},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={arXiv:COMING\_SOON},
}