1.Current Status, Trends, and Opportunities in the Study of Computable Phenotypes for Rare Diseases
Jindong WU ; Qiaorui WEN ; Jian GUO ; Shengfeng WANG
JOURNAL OF RARE DISEASES 2026;5(1):90-99
Disease computable phenotype is a data model designed to identify specific clinical conditions or characteristics, which automatically extracts information from clinical databases such as electronic health records through algorithms. Phenotypic data for rare diseases often reside in unstructured text. Due to the scarcity of rare disease cases, atypical symptoms, and insufficient physician experience, misdiagnosis and underdiagnosis rates remain high. In this context, the application of computable phenotype technology holds promise for improving the accuracy and efficiency of rare disease diagnosis. This article reviews the current research status, challenges, and opportunities of computable phenotype technology in biomedicine, particularly in the field of rare diseases, and proposes a development and validation framework for rare disease computable phenotypes, aiming to provide research and development insights for computable phenotypes to empower the diagnosis and treatment of rare diseases.
2.Sleep modes based on objective measurement and diseases of the body systems:a cohort study of 87 617 participants from the UK Biobank dataset
Yimeng WANG ; Qing CHEN ; Siwen LUO ; Fuquan SHI ; Mengchao HE ; Shengfeng WANG ; Qiaorui WEN ; Yingzhong DAI ; Hao QU ; Jia CAO
Journal of Army Medical University 2025;47(4):318-325
Objective To investigate the impact of sleep modes on the risk for diseases of the body systems.Methods Based on a subset of UK Biobank dataset comprising 87 617 participants,3 sleep dimensions including 6 sleep indicators were obtained through a wrist-worn accelerometer,that is sleep duration and onset,sleep rhythm(relative amplitude and stability),and sleep quality(sleep efficiency and number of awakenings).Latent profile analysis(LPA)was applied to identify and classify distinct sleep modes.Then their longitudinal medical records were the association between different sleep modes and the risk for 467 diseases.Results LPA identified 5 subgroups of unique sleep modes in the participants.Among the 5 subgroups,the subgroup 4 had relatively optimal levels in above sleep indicators.Compared to the subgroup 4,the other 4 subgroups exhibited variations in different sleep dimensions,with at least one indicator demonstrating an unfavorable trend.These subgroups also revealed differences in racial composition,shift work and social deprivation index.Moreover,there were notable differences in the risk of various system diseases among the subgroups(P<0.05).When compared to the subgroup 4,the other 4 subgroups exhibited an elevated risk for certain diseases(comprising a total of 126 diseases),with the diseases of the circulatory system,digestive system and musculoskeletal system most common.Among the 5 subgroups,the subgroup 2(shorter sleep duration and later sleep onset)and the subgroup 5(rhythm disorder)had the highest counts of associated risks,amounting to 85 and 91 types,respectively,but there was certain difference in their systematic composition.Conclusion There are different sleep modes within the participants,and the modes are potentially associated with an increased risk for diseases of body systems.Comprehensive interventions targeting overall sleep modes rather than single sleep indicator may yield obvious health benefits.
3.Distribution characteristics of smoking behavior among adult twins in China
Shunkai LIU ; Wenjing GAO ; Weihua CAO ; Jun LYU ; Canqing YU ; Shengfeng WANG ; Tao HUANG ; Dianjianyi SUN ; Chunxiao LIAO ; Yuanjie PANG ; Ruqin GAO ; Min YU ; Jinyi ZHOU ; Xianping WU ; Zhong DONG ; Fan WU ; Dezheng WANG ; Zhihua XU ; Yu LIU ; Jianrui WANG ; Jie YIN ; Shengli YIN ; Liming LI
Chinese Journal of Preventive Medicine 2025;59(7):1090-1096
This study aims to describe the population and regional distribution characteristics of smoking behavior among adult twins in the China Twin Registry (CNTR), as well as the concordance rates for smoking behavior in monozygotic and dizygotic twins, and estimate the heritability. The study population included adult twins in CNTR who had smoking questionnaire data. A random-effects regression model was used to describe the distribution of smoking behavior among different subgroups based on various characteristics. The concordance of smoking behavior between different zygosity groups was calculated, and heritability was estimated. A total of 28 444 twin pairs were included in this study, with an average age of (36.6±12.0) years. Among male twins, 41.2% were current smokers, while only 1.2% of females smoked. Higher smoking rates were observed among male smokers in the 50-59 age group ( z=23.0, P<0.001), northern regions ( z=2.9, P<0.01), rural areas ( z=-5.2, P<0.001), those who were divorced/widowed ( z=3.8, P<0.001), and first-born twins ( z=-4.3, P<0.001), while lower smoking rates were found in those with higher education ( z=-16.1, P<0.001) and unmarried individuals ( z=-16.0, P<0.001). The smoking concordance rate for male monozygotic twins was 69.6%, significantly higher than the 57.3% concordance rate for dizygotic twins ( χ 2=105.0, P<0.05). The heritability of smoking behavior in male twins was estimated at 28.9% (95% CI: 24.3%-33.4%). Stratified analyses showed differences in heritability across regions and age groups: the heritability in northern regions was 32.6% (95% CI: 27.3%-38.0%), higher than the 21.0% (95% CI: 12.4%-29.5%) observed in southern regions; the highest heritability of 35.1% (95% CI: 26.3%-43.9%) was found in the 18-29 age group, with heritability decreasing with age. In conclusion, the smoking rate and influencing factors in the twin population are similar to those in the general population, with unique characteristics, such as higher smoking rates in first-born twins. Genetic factors have a significant impact on smoking behavior.
4.Incidence and influencing factors of ocular surface disease among power grid construction workers in plateau: a real-world study
Xinyu YANG ; Yunjing ZHANG ; Huziwei ZHOU ; Quanquan GONG ; Xinyu WANG ; Xiaoyu ZHANG ; Zhixia LI ; Shiming LI ; Shengfeng WANG
Chinese Journal of Experimental Ophthalmology 2025;43(5):443-451
Objective:To analyze the incidence and risk factors of ocular surface disease among power grid construction workers in plateau.Methods:A total of 11 132 construction personnel from the Ngari prefecture-central Tibet power grid interconnection project were included from 2019 to 2020.Baseline characteristics including age, gender, body mass index, developmental and nutritional status, relevant clinical indicators, etc.and follow-up data regarding incidence of ocular surface diseases were obtained from the medical records of Ali interconnection project staff medical station.The altitude of workplace and residence of the study population were obtained from the website (https: //zh-cn.topographic-map.com/legal/).The mean age of the subjects was (36.17±10.48) years, of which 95.33%(10, 612 subjects) were male.The median follow-up time was 1.53 years.The altitude of the residence and workplace were (1 954.77±940.64) and (4 535.09±232.71) meters, respectively.The incidence of ocular surface diseases in groups with different characteristics was calculated.Differential variables for the incidence of ocular surface diseases were screened by univariate Cox proportional hazards regression model.Influencing factors of ocular surface diseases multivariate were explored by Cox proportional hazards model.This study was approved by the Ethics Committee of Peking University Health Science Center (No.IRB00001052-21066).Results:During the follow-up period, the incidence of ocular surface disease was 9.27% (1 032 cases), and the incidence of conjunctivitis and keratitis was 6.58% (733 cases) and 1.80% (200 cases), respectively.Multivariate Cox proportional hazards regression analysis showed that for every 1 000 meters increase in altitude of residence, the risk of ocular surface disease decreased by 15% ( HR[95% CI]: 0.85[0.80~0.91], P<0.001).For every 100 meters increase in altitude of workplace, the risk of ocular surface disease increased by 5% ( HR[95% CI]: 1.04[1.01~1.07], P=0.006).Decreased blood oxygen saturation ( HR[95% CI]: 1.09[1.02~1.16], P=0.007), hearing pulmonary dry rales (hazard ratio ( HR)[95% CI]: 1.53[1.12~2.09], P=0.007) and heart murmurs ( HR[95% CI]: 4.44[1.43~13.83], P=0.010) were associated with ocular surface disease. Conclusions:The incidence of ocular surface disease in personnel engaged in electric grid construction at high altitudes should not be ignored.High working altitude, low residence altitude, pulmonary dry rales, heart murmurs and low blood oxygen saturation are factors associated with the incidence of ocular surface disease.
5.Insights on facilitators and barriers to regulating non-medical use of prescription opioids:a qualitative study
Yuehan DUAN ; Huziwei ZHOU ; Yingzi YANG ; Qiaorui WEN ; Hongling CHU ; Jingling WANG ; Zhiqin JIANG ; Yexiang SUN ; Yu ZHU ; Shengfeng WANG
Chinese Journal of Pharmacoepidemiology 2025;34(11):1265-1275
Objective The aim is to understand the common scenarios of non-medical use of prescription opioids(NMUPO)and analyze the potential facilitating and hindering factors in the regulatory process of NMUPO from the perspective of healthcare professionals.Methods Healthcare professionals in local hospitals were surveyed through a two-stage purposive sampling from June to August 2022 in Ningbo,China.The survey was conducted using a semi-structured questionnaire on topics,and thematic analysis were used to identify and summarise key themes and patterns.Results A total of 75 participants were included,the average age was(43.9±7.2)years,and 54(72.0%)were male.The most common NMUPO scenarios involved middle-aged males pretending acute severe pain to obtain injectable opioids.The facilitating and hindering factors related to the regulation of NMUPO can be categorized into three types:institutional governance,technical support,and individual behaviors.At the institutional level,facilitating factors included strict national prescribing policies and local"narcotic drug card"systems,while barriers comprised incomplete lists of controlled substances.At the technological support level,facilitating factors included the establishment of regional health information platforms,while barriers included the lack of standardized prescription guidelines and diagnostic decision-support tools.At the individual level,facilitating factors included the public's cautious attitude toward drug misuse,while barriers included strained doctor-patient relationships.Conclusion China still faces significant challenges in addressing NMUPO and urgently needs to improve the existing regulatory system.It is recommended that reforms be carried out in areas such as pharmaceutical control mechanisms,drug treatment and rehabilitation services,preventive health education activities,and the optimized use of health information systems.
6.Development and application of a rapid identification algorithm for cutaneous lupus erythematosus and its subtypes based on medical insurance databases
Yutong WANG ; Xianglong MENG ; Yu PAN ; Chen WEI ; Hui JIN ; Shengfeng WANG
Chinese Journal of Pharmacoepidemiology 2025;34(7):743-752
Objective To develop and validate data extraction and patient identification algorithms for cutaneous lupus erythematosus(CLE)and its two subtypes,discoid lupus erythematosus(DLE)and subacute cutaneous lupus erythematosus(SCLE),and to enable high-efficiency patient identification in large-scale electronic health databases.Methods This study utilized data from the 2013-2017 National Insurance Claims for Epidemiological Research(NICER)to construct data extraction and rapid patient identification algorithms.The manual verification results were used as gold standard to assess the sensitivity and specificity of the algorithms.Additionally,the basic characteristics of the identified patients were analyzed.Results Initially,standardized expressions were developed based on medical terminology and diagnostic codes.These were further refined with input from clinicians to include potential synonyms and common misspellings,improving the preliminary screening expressions.Through iterative verification by clinicians and data management engineers,a final disease-specific screening algorithm was established.The developed extraction and identification algorithms for all 3 targeted disease demonstrated strong performance,with sensitivity values of 0.985,1.000,and 0.991,and specificity values of 0.997,0.999,and 0.998 for CLE,DLE,and SCLE,respectively.A total of 34,554 CLE cases,including 2,879 DLE cases,and 623 SCLE cases were identified between 2013 and 2017,with a higher prevalence among females than males.Conclusion This study developed and validated an identification algorithm for CLE patients based on medical insurance databases,demonstrating high performance.The proposed algorithm provides a methodological framework and empirical evidence for designing and optimizing big data-driven rapid patient identification algorithms in dermatology research.
7.Expert consensus for off-label drug use of rare disease:a protocol
Chaoyang CHEN ; Yuehan DUAN ; Lin ZHUO ; Guohua HE ; Yanqin ZHANG ; Ying ZHOU ; Shengfeng WANG ; Yimin CUI ; Jie DING
Chinese Journal of Pharmacoepidemiology 2025;34(9):1066-1073
Rare diseases are a collective term for diseases with extremely low prevalence and incidence rates.Up to now,China has released two lists identifying a total of 207 rare diseases.Given that most rare diseases do not have drugs with corresponding indications,physicians frequently resort to using off-label drugs when treating patients with rare diseases.However,there is currently no systematic guideline or expert consensus for the use of off-label medications in China.To comprehensively collect existing evidence of off-label drug use for rare diseases,fully analyze and evaluate the rationality of off-label drug use for rare diseases,and standardize the management of off-label drug use for rare diseases,the Rare Disease Branch of Beijing Medical Association,Chinese Pharmaceutical Association,Beijing Pharmaceutical Association,and the School of Public Health,Peking University have jointly initiated the drafting of the Expert Consensus on Off-label Use of Drugs for Rare Diseases.This consensus refer to the WHO Handbook for Guideline Development,the Guidelines for Developing/Revising Clinical Diagnostic and Treatment Guidelines in China(2022 Edition),the AGREE Ⅱ and the STAR tools.This protocol outlines the background and purpose of consensus,as well as the comprehensive framework for consensus development,encompassing panel formation,clinical issue identification,evidence retrieval,data extraction,and evidence-based recommendation formulation.
8.Large language models empowering pharmacoepidemiology research
Shucheng SI ; Liuliu WU ; Conghui WANG ; Ziming YANG ; Jian DU ; Shengfeng WANG ; Siyan ZHAN
Chinese Journal of Pharmacoepidemiology 2025;34(9):1074-1083
The emergence of artificial intelligence(AI)has had a significant impact on medical research and practice,both in terms of the number of studies and research paradigms,and has become an important tool for the development of pharmacoepidemiology.However,traditional AI has faced many challenges,while facilitating pharmacoepidemiology research,such as complex data processing,difficulty in identifying drug exposures and potential outcomes,and time-consuming and laborious study design and implementation.The rapid development of generative AI,represented by large language models(LLMs),has demonstrated a unique potential to enhance research efficiency,shift research paradigms,and facilitate knowledge discovery.LLMs are equipped with natural language understanding and generation capabilities.Through deep mining of multi-dimensional data resources,LLMs can quickly and accurately extract,analyze,summarize,and present the required information,which can not only help drug discovery,drug repurposing,pharmacovigilance and other pharmacoepidemiological tasks,but also provide powerful support for the whole process of research protocol design,data analysis,result interpretation and paper publication.Driven by LLMs,pharmacoepidemiology research is gradually moving into a new stage based on big data and automated analysis.Of course,LLMs also have problems of data bias,"illusion"of results,and ethical and legal regulation.By strengthening interdisciplinary cooperation,establishing a standardized evaluation system,improving ethical and regulatory guidance,enhancing data quality,strengthening practitioner training and capacity building,and promoting human-machine collaborative research modes,it is expected that the potential of LLMs in pharmacoepidemiology will be fully released,and it will provide a more scientific,rapid,and efficient technological support for drug regulation and public health decision-making.
9.Current approaches and challenges in addressing class imbalance in medical prediction models
Xianglong MENG ; Yutong WANG ; Xin ZHANG ; Siyan ZHAN ; Shengfeng WANG
Chinese Journal of Epidemiology 2025;46(9):1632-1639
With the rise of personalized medicine and the rapid development of big data technology, medical prediction models have become increasingly important in disease diagnosis, prognosis assessment, and risk stratification. However, class imbalance is a common problem in medical data, which can result in models being overly trained toward the majority class rather than the minority class, influencing the detection power and clinical application value. This paper systematically summarizes traditional methods in addressing class imbalance, including data pre-processing and algorithm level strategies, and introduces the applications of new technologies such as generative adversarial networks and transfer learning and suggests key considerations and potential research focus for addressing class imbalance to provide reference for researchers to select appropriate strategies.
10.Artificial intelligence in epidemiology: a decade-long bibliometric analysis
Conghui WANG ; Ziming YANG ; Wei SHI ; Chengwei XI ; Shucheng SI ; Liuliu WU ; Jian DU ; Shengfeng WANG ; Siyan ZHAN
Chinese Journal of Epidemiology 2025;46(9):1650-1659
Objective:To describe the hotspots and application trends of artificial intelligence (AI) in epidemiology in the past decade and analyze its advantages and challenges.Methods:The literatures with AI and epidemiology related keywords were systematically retrieved from Web of Science and China National Knowledge Infrastructure from 2014 to 2024. CiteSpace was used for bibliometric analysis of publication volume, keyword co-occurrence, clustering, emergence and cited literature co-occurrence analysis.Results:A total of 5 389 English papers and 1 659 Chinese papers were included, showing an increasing publication trend. High-frequency Chinese keywords included prediction, influencing factor, and machine learning, while English keywords frequently used were machine learning, prediction, and artificial intelligence. The Chinese keywords formed 14 clusters such as epidemiological characteristic, dietary pattern, and elderly individual, and the English keywords formed 21 clusters including prediction model, risk factor, and adult. In international studies, health policy, COVID-19, and digital health were the emerging frontier keywords. Eleven core papers were selected, covering key areas like traffic accident risk assessment, public health big data application, and deep learning in medical diagnosis.Conclusions:This study systematically summarized the research hotspots and development trends of AI applications in epidemiology over the past decade by using bibliometric methods, which indicated that current AI-based epidemiological studies are still in the exploratory phase, with the coexisting of both advantages and challenges. Continued attention should be paid to the future development of this field.

Result Analysis
Print
Save
E-mail