Abstract
Medicine is fundamentally an empirical body of knowledge accumulated through long-term observation, validation, and refinement, and ultimately realized in the messy, high-variance real clinical practice. Physicians’ diagnostic-and-therapeutic competence is gradually forged through repeated cycles of “application–reflection–improvement,” crystallizing into distinctive, individualized methodologies. Considering the fact of the substantial variation in treatment outcomes, the formation of master physicians’ knowledge systems is necessary but time-consuming and their transmission is often limited in scope, making high-quality expertise difficult to scale and disseminate, and the scarcity of advanced clinical resources. To mitigate these challenges, we propose Med-Shicheng, a general framework that enables large language model to systematically learn and transfer distinguished physicians’ diagnostic-and-therapeutic philosophy, and case-dependent adaptation rules in a standardized manner. Built upon Tianyi, Med-Shicheng comprises five elaborately designed five stages. We take five National Masters of Chinese Medicine/distinguished TCM physicians as representative targets, curate and organize multi-source materials and train a single model to simultaneously internalize the five distinctive knowledge systems across seven tasks: etiology–pathogenesis analysis, syndrome diagnosis, determination of treatment principles, prescription generation, prescription explanation, symptom evolution with corresponding regimen modifications, and clinical advice. Implemented on Qwen2.5-1.5B-Base, Med-Shicheng can be deployed on resource-constrained GPUs, while demonstrating evaluation performance comparable to DeepSeek-R1 and GPT-5. Besides, we further analyze the reliability of LLM-as-a-Judge against physician-based assessment. It shows that although automated judging captures broad performance trends, it exhibits noticeable biases in fine-grained, individualized clinical distinctions, indicating that physician involvement remains necessary when ground truth is unavailable and that judge models require targeted medical domain adaptation to achieve reliable evaluation.
Med-Schiehng TCM heritage Framework
We propose Med-Shicheng, a general paradigm that enables lightweight LLMs to inherit the diagnostic-and-therapeutic (D&T) expertise of distinguished TCM doctors. The framework contains five progressive stages: (1) building a medical in-domain foundation model, (2) learning regular TCM D&T strategies, (3) reasoning-enhanced D&T fine-tuning, (4) fine-tuning on each doctor’s unique medical knowledge, and (5) reinforcement-style refinement of their distinctive reasoning and treatment strategies. For every stage, we curate task-specific datasets that follow TCM clinical logic and cover the full chain from etiology analysis to prescription and follow-up adjustment. This staged design turns a general-purpose LLM into a doctor-aware medical agent that can internalize multiple masters’ styles within a single model and switch behaviors according to the target physician.For more detailed information, please see our paper.
Evaluatino Benchmark
Our evaluation framework combines automated LLM-as-a-judge scoring with expert TCM assessment. For automatic evaluation, we use state-of-the-art general LLMs, DeepSeek-V3.2 and GPT-5, as judges. Each model’s response to a clinical case is scored twice, once by each judge, who compare the generated answer with the gold label and decide how clinically desirable it is. Because every task response is a long, structured text spanning seven components, classical token-level metrics such as accuracy, recall or F1-score cannot reflect quality; holistic judging by strong LLMs is more suitable. To complement this, we conduct human evaluation with 18 senior TCM physicians, grouped around the five target masters they directly studied under. Using Delphi consensus, we provide unified criteria and ask them to rate all models along five dimensions: (1) similarity to the target master’s D&T style, (2) consistency with TCM philosophical principles, (3) safety and risk control, (4) therapeutic completeness, and (5) clinical coherence.
Automatic Evaluations
Automatic evaluation by DeepSeek-V3.2 and GPT-5 shows broadly consistent trends across all target TCM masters. One notable exception is HuatuoGPT2-7B on Dr. Huaitang Du’s cases, where it achieves unusually high scores for prescriptions and treatment principles by reproducing label-like content, suggesting possible training-data overlap. Overall, models naturally split into two tiers. The high-performing group includes Med-Shicheng, GPT-5, DeepSeek-R1, Qwen3-235B-A22B-Thinking, and Gemini-2.5-Pro; the lower-performing group includes Qwen2.5-1.5B-Instruct, Tianyi, and HuatuoGPT2. Med-Shicheng is the only lightweight model in the top tier, delivering content quality comparable to hundred-billion-parameter systems. However, no single model dominates every doctor or metric. Using DeepSeek-V3.2 as judge, DeepSeek and GPT-5 often rank in the top two, but Med-Shicheng, Gemini-2.5-Pro and Qwen3-235B-A22B-Thinking sometimes outperform them, for example on the cases of Fengchun Wang, Shaoqin Zhao and Huaitang Du.
Human Doctors Evaluations
Human evaluation reveals a different picture from LLM-as-a-judge scoring. We ask 18 senior TCM physicians to rate all models on each master’s cases, along five dimensions: similarity to the target physician, adherence to TCM theory, safety, therapeutic completeness, and clinical coherence. Overall, human doctors assign Med-Shicheng higher and more stable scores for most masters, including Huaitang Du, Xiaohong Gu, Fengchun Wang and Bowei Qin, whereas GPT-5 and DeepSeek-V3.2 do not consistently place it near the top. This suggests that very large general LLMs can judge long, complex answers under well-designed prompts, but struggle with fine-grained distinctions in individualized clinical style. They tend to evaluate with generic TCM knowledge, even for widely documented experts such as Dr. Du, while human physicians can recognize subtle signatures of each master’s system. The gap is especially visible for Bowei Qin’s cases, where clinicians rank Med-Shicheng near second, but LLM judges push it down to around fourth.
BibTeX
@article{medshicheng2026,
title={Med-Shicheng:From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLM},
author={First Author and Second Author and Third Author},
year={2026},
eprint={COMING SOON},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={arXiv:COMING\_SOON},
}