1.A Machine Learning Approach to Reference Interval Estimation for Red Cell Parameters in a South and East Asian Population
Veera Sekaran NADARAJAN ; Pavai STHANESHWAR ; Jia Qi LIM ; Angeli AMBAYYA ; Putri Junaidah Megat YUNUS
Annals of Laboratory Medicine 2026;46(1):41-51
Background:
Iron deficiency (ID) and hemoglobinopathies are highly prevalent in Southeast Asia. Accurate estimation of reference intervals (RIs) for red cell parameters is complicated by the need to exclude individuals with these conditions from the reference population. Indirect RI estimations using machine learning could help overcome these challenges.
Methods:
We developed a binary classification model using eXtreme Gradient Boosting (XGB) to distinguish normal individuals from those with ID, hemoglobinopathies, or other anemias. The model was trained on an annotated dataset comprising 5,520 complete blood count (CBC) results and validated with a holdout dataset of 2,367 CBC results. An independent dataset of 64,100 CBC results was used to identify individuals predicted to be normal, from which RIs were estimated using the refineR algorithm.
Results:
The XGB model achieved an area under the ROC of 0.97 (95% confidence interval: 0.96–0.97) for distinguishing between individuals with normal versus abnormal values. Among individuals within the independent dataset, 40,300 (62.9%) were predicted to be normal. The refineR-based reference limits (RLs) derived from this subset approximated those obtained through a direct approach. Improvements in the accuracy of indirect RL estimates were most evident for hematocrit, hemoglobin, and red cell concentrations.
Conclusions
Combining XGB with refineR to indirectly derive RIs for red cell parameters improved the accuracy and yielded results comparable with those of directly derived RIs. A further benefit was the capacity to generate sex- and age-specific ranges, which has remained difficult to achieve through direct approaches.
2.Reference interval establishment of full blood count extended research parameters in the multi-ethnic population of Malaysia
Angeli Ambayya ; Andrew Octavian Sasmita ; Qian Yun Zhang ; Anselm Su Ting ; Chang Kian Meng ; Jameela Sathar ; Subramanian Yegappan
The Medical Journal of Malaysia 2019;74(6):534-536
Haematological cellular structures may be elucidated using
automated full blood count (FBC) analysers such as Unicel
DxH 800 via cell population data (CPD) analysis. The CPD
values are generated by calculating volume, conductivity,
and five types of scatter angles of individual cells which
would form clusters or populations. This study considered
126 CPD parameter values of 1077 healthy Malaysian adults
to develop reference intervals for each CPD parameter. The
utility of the CPD reference interval established may range
from understanding the normal haematological cellular
structures to analysis of distinct cellular features related to
the development of haematological disorders and
malignancies.
3.Microarray-Based Genomic Analysis Identifies Germline and Somatic Copy Number Variants and Loss of Heterozygosity in Acute Myeloid Leukaemia
Malaysian Journal of Medicine and Health Sciences 2018;14(SP3):11-24
Introduction: Insights into molecular karyotyping using comparative genomic hybridization (CGH) and single nucleotide polymorphism (SNP) arrays enable the identification of copy number variations (CNVs) at a higher resolution and facilitate the detection of copy neutral loss of heterozygosity (CN-LOH) otherwise undetectable by conventional cytogenetics. The applicability of a customised CGH+SNP 180K DNA microarray in the diagnostic evaluation of Acute Myeloid Leukaemia (AML) in comparison with conventional karyotyping was assessed in this study. Methods: Paired tumour and germline post induction (remission sample obtained from the same patient after induction) DNA were used to delineate germline variants in 41 AML samples and compared with the karyotype findings. Results: After comparing the tumour versus germline DNA, a total of 55 imbalances (n 5-10 MB = 21, n 10-20 MB = 8 and n >20 MB = 26) were identified. Gains were most common in chromosome 4 (26.7%) whereas losses were most frequent in chromosome 7 (28.6%) and X (25.0%). CN-LOH was mostly seen in chromosome 4 (75.0%). Comparison between array CGH+SNP and karyotyping revealed 20 cases were in excellent agreement and 13 cases did not concord whereas in 15 cases finding could not be confirmed as no karyotypes available. Conclusion: The use of a combined array CGH+SNP in this study enabled the detection of somatic and germline CNVs and CN-LOHs in AML. Array CGH+SNP accurately determined chromosomal breakpoints compared to conventional cytogenetics in relation to presence of CNVs and CN-LOHs.
Acute myeloid leukaemia


Result Analysis
Print
Save
E-mail