The importance of disease incidence rate on performance of GBLUP, threshold BayesA and machine learning methods in original and imputed data set
Abstract
Aim of study: To predict genomic accuracy of binary traits considering different rates of disease incidence.
Area of study: Simulation
Material and methods: Two machine learning algorithms including Boosting and Random Forest (RF) as well as threshold BayesA (TBA) and genomic BLUP (GBLUP) were employed. The predictive ability methods were evaluated for different genomic architectures using imputed (i.e. 2.5K, 12.5K and 25K panels) and their original 50K genotypes. We evaluated the three strategies with different rates of disease incidence (including 16%, 50% and 84% threshold points) and their effects on genomic prediction accuracy.
Main results: Genotype imputation performed poorly to estimate the predictive ability of GBLUP, RF, Boosting and TBA methods when using the low-density single nucleotide polymorphisms (SNPs) chip in low linkage disequilibrium (LD) scenarios. The highest predictive ability, when the rate of disease incidence into the training set was 16%, belonged to GBLUP, RF, Boosting and TBA methods. Across different genomic architectures, the Boosting method performed better than TBA, GBLUP and RF methods for all scenarios and proportions of the marker sets imputed. Regarding the changes, the RF resulted in a further reduction compared to Boosting, TBA and GBLUP, especially when the applied data set contained 2.5K panels of the imputed genotypes.
Research highlights: Generally, considering high sensitivity of methods to imputation errors, the application of imputed genotypes using RF method should be carefully evaluated.Downloads
References
Bishop CM, 2006. Pattern recognition and machine learning (information science and statistics). Springer-Verlag, NY.
Bohlouli M, Alijani S, Javaremi AN, König S, Yin T, 2017. Genomic prediction by considering genotype × environment interaction using different genomic architectures. Ann Anim Sci 17: 683-701. https://doi.org/10.1515/aoas-2016-0086
Breiman L, 2001. Random forests. Machine Learning 45: 5-32. https://doi.org/10.1023/A:1010933404324
Chen L, Li C, Sargolzaei M, Schenkel F, 2014. Impact of genotype imputation on the performance of GBLUP and Bayesian methods for genomic prediction. PLoS One 9: e101544. https://doi.org/10.1371/journal.pone.0101544
Daetwyler HD, Calus MP, Pong-Wong R, de los Campos G, Hickey JM, 2013. Genomic prediction in animals and plants: simulation of data, validation, reporting, and benchmarking. Genetics 193: 347-365. https://doi.org/10.1534/genetics.112.147983
De Los Campos G, Naya H, Gianola D, Crossa J, Legarra A, Manfredi E, Weigel K, Cotes JM, 2009. Predicting quantitative traits with regression models for dense molecular markers and pedigree. Genetics 182: 375-385. https://doi.org/10.1534/genetics.109.101501
Egger-Danner C, Cole J, Pryce J, Gengler N, Heringstad B, Bradley A, Stock KF, 2015. Invited review: overview of new traits and phenotyping strategies in dairy cattle with a focus on functional traits. Animal 9: 191-207. https://doi.org/10.1017/S1751731114002614
Felipe VP, Okut H, Gianola D, Silva MA, Rosa GJ, 2014. Effect of genotype imputation on genome-enabled prediction of complex traits: an empirical study with mice data. BMC Gent 15: 149. https://doi.org/10.1186/s12863-014-0149-9
Freund Y, Schapire RE, 1996. Experiments with a new boosting algorithm. Icml 96: 148-156. https://dl.acm.org/doi/10.5555/3091696.3091
Friedrich J, Antolín R, Edwards S, Sánchez‐Molano E, Haskell M, Hickey J, Wiener P, 2018. Accuracy of genotype imputation in Labrador Retrievers. Anim Genet 49: 303-311. https://doi.org/10.1111/age.12677
Ghafouri-Kesbi F, Rahimi-Mianji G, Honarvar M, Nejati-Javaremi A, 2017. Predictive ability of Random Forests, Boosting, Support Vector Machines and Genomic Best Linear Unbiased Prediction in different scenarios of genomic evaluation. Anim Prod Sci 57: 229-236. https://doi.org/10.1071/AN15538
Goddard M, 2009. Genomic selection: prediction of accuracy and maximisation of long term response. Genetica 136: 245-257. https://doi.org/10.1007/s10709-008-9308-0
González-Recio O, Forni S, 2011. Genome-wide prediction of discrete traits using Bayesian regressions and machine learning. Genet Sel Evol 43: 7. https://doi.org/10.1186/1297-9686-43-7
Guo Z, Tucker DM, Basten CJ, Gandhi H, Ersoz E, Guo B, Xu Z, Wang D, Gay G, 2014. The impact of population structure on genomic prediction in stratified populations. Theor Appl Genet 127: 749-762. https://doi.org/10.1007/s00122-013-2255-x
Hayes B, Daetwyler H, Bowman P, Moser G, Tier B, Crump R, Khatkar M, Raadsma H, Goddard M, 2009. Accuracy of genomic selection: comparing theory and results. Proc Assoc Advmt Anim Breed Genet, pp: 34-37.
Hickey JM, Crossa J, Babu R, de los Campos G, 2012. Factors affecting the accuracy of genotype imputation in populations from several maize breeding programs. Crop Sci 52: 654-663. https://doi.org/10.2135/cropsci2011.07.0358
Kabisch M, Hamann U, Bermejo JL, 2017. Imputation of missing genotypes within LD-blocks relying on the basic coalescent and beyond: consideration of population growth and structure. BMC genomics 18: 798. https://doi.org/10.1186/s12864-017-4208-2
Lakhssassi K, González-Recio O, 2017. A haplotype regression approach for genetic evaluation using sequences from the 1000 bull genomes Project. Span J Agric Res 15 (4): e0407. https://doi.org/10.5424/sjar/2017154-11736
Liu H, Zhou H, Wu Y, Li X, Zhao J, Zuo T, Zhang X, Zhang Y, Liu S, Shen Y, 2015. The impact of genetic relationship and linkage disequilibrium on genomic selection. PLoS One 10: e0132379. https://doi.org/10.1371/journal.pone.0132379
Madsen P, Jensen J, 2013. A users guide to DMU. A package for analysing multivariate mixed models, Version 6. Center for Quantitative Genetics and Genomics, University of Aarhus, Denmark. https://dmu.ghpc.au.dk/
Mc Hugh N, Meuwissen T, Cromie A, Sonesson A, 2011. Use of female information in dairy cattle genomic breeding programs. J Dairy Sci 94: 4109-4118. https://doi.org/10.3168/jds.2010-4016
Meuwissen T, Hayes B, Goddard M, 2001. Prediction of total genetic value using genome-wide dense marker maps. Genetics 157: 1819-1829.
Naderi S, Yin T, König S, 2016. Random forest estimation of genomic breeding values for disease susceptibility over different disease incidences and genomic architectures in simulated cow calibration groups. J Dairy Sci 99: 7261-7273. https://doi.org/10.3168/jds.2016-10887
Naderi S, Bohlouli M, Yin T, König S, 2018. Genomic breeding values, SNP effects and gene identification for disease traits in cow training sets. Anim Genet 49: 178-192. https://doi.org/10.1111/age.12661
Naderi Y, Sadeghi S, 2019. Assessment of the genomic prediction accuracy of discrete traits with imputation of missing genotypes. Anim Sci Papers Rep 37: 149-168.
Pausch H, MacLeod IM, Fries R, Emmerling R, Bowman PJ, Daetwyler HD, Goddard ME, 2017. Evaluation of the accuracy of imputed sequence variant genotypes and their utility for causal variant detection in cattle. Genet Sel Evol 49: 24. https://doi.org/10.1186/s12711-017-0301-x
Pimentel EC, Wensch-Dorendorf M, König S, Swalve HH, 2013. Enlarging a training set for genomic selection by imputation of un-genotyped animals in populations of varying genetic architecture. Genet Sel Evol 45: 12. https://doi.org/10.1186/1297-9686-45-12
Sadeghi S, Rafat sA, Alijani S, 2018. Evaluation of imputed genomic data in discrete traits using Random forest and Bayesian threshold methods. Acta Sci Anim Sci 40: e39007. https://doi.org/10.4025/actascianimsci.v40i1.39007
Sargolzaei M, Schenkel FS, 2009. QMSim: a large-scale genome simulator for livestock. Bioinformatics 25: 680-681. https://doi.org/10.1093/bioinformatics/btp045
Sargolzaei M, Chesnais J, Schenkel F, 2011. FImpute-An efficient imputation algorithm for dairy cattle populations. J Dairy Sci 94: 421.
Su G, Madsen P, 2013. User's Guide for GMATRIX version 2, a program for computing genomic relationship matrix.
VanRaden PM, 2008. Efficient methods to compute genomic predictions. J Dairy Sci 91: 4414-4423. https://doi.org/10.3168/jds.2007-0980
Ventura RV, Miller SP, Dodds KG, Auvray B, Lee M, Bixley M, Clarke SM, McEwan JC, 2016. Assessing accuracy of imputation using different SNP panel densities in a multi-breed sheep population. Genet Sel Evol 48: 71. https://doi.org/10.1186/s12711-016-0244-7
Wang C, Ding X, Wang J, Liu J, Fu W, Zhang Z, Yin Z, Zhang Q, 2013. Bayesian methods for estimating GEBVs of threshold traits. Heredity 110: 213-219. https://doi.org/10.1038/hdy.2012.65
Wang Y, Lin G, Li C, Stothard P, 2016. Genotype imputation methods and their effects on genomic predictions in cattle. Spr Sci Rev 4: 79-98. https://doi.org/10.1007/s40362-017-0041-x
Wang C, Li X, Qian R, Su G, Zhang Q, Ding X, 2017. Bayesian methods for jointly estimating genomic breeding values of one continuous and one threshold trait. PloS One 12: e0175448. https://doi.org/10.1371/journal.pone.0175448
Wang Q, Yu Y, Yuan J, Zhang X, Huang H, Li F, Xiang J, 2017. Effects of marker density and population structure on the genomic prediction accuracy for growth trait in Pacific white shrimp Litopenaeus vannamei. BMC Gent 18: 45. https://doi.org/10.1186/s12863-017-0507-5
Wientjes YC, Calus MP, Goddard ME, Hayes BJ, 2015. Impact of QTL properties on the accuracy of multi-breed genomic prediction. Genet Sel Evol 47: 42. https://doi.org/10.1186/s12711-015-0124-6
Wimmer V, Auinger HJ, Albrecht T, Schoen CC, 2015. Framework for the analysis of genomic prediction data using R (synbreed). https://cran.rproject.org/web/packages/synbreed/index.html
Yang P, Hwa Yang Y, B Zhou B, Y Zomaya A, 2010. A review of ensemble methods in bioinformatics. Curr Bioinform 5: 296-308. https://doi.org/10.2174/157489310794072508
Yin T, Pimentel E, Borstel UKv, König S, 2014. Strategy for the simulation and analysis of longitudinal phenotypic and genomic data in the context of a temperature× humidity-dependent covariate. J Dairy Sci 97: 2444-2454. https://doi.org/10.3168/jds.2013-7143
© CSIC. Manuscripts published in both the print and online versions of this journal are the property of the Consejo Superior de Investigaciones Científicas, and quoting this source is a requirement for any partial or full reproduction.
All contents of this electronic edition, except where otherwise noted, are distributed under a Creative Commons Attribution 4.0 International (CC BY 4.0) licence. You may read the basic information and the legal text of the licence. The indication of the CC BY 4.0 licence must be expressly stated in this way when necessary.
Self-archiving in repositories, personal webpages or similar, of any version other than the final version of the work produced by the publisher, is not allowed.









