A Tutorial on Linear Classification Methods for Educational and Behavioral Research with R

Authors

  • Farzan Madadizadeh Medical Informatics Research Center, Institute for Futures Studies in Health, Kerman University of Medical Sciences, Kerman, Iran and Center for healthcare Data modeling, Departments of Biostatistics and Epidemiology, School of public health, Shahid Sadoughi University of Medical Sciences, Yazd, Iran
  • Rashed Pourhamidi PhD Student, Departments of Biostatistics and Epidemiology, School of public health, Shahid Sadoughi University of Medical Sciences, Yazd, Iran and Non Communicable Diseases Research Center, Bam University of Medical Sciences, Bam, Iran
  • Nahid khoddami Center for healthcare Data modeling, Departments of Biostatistics and Epidemiology, School of public health, Shahid Sadoughi University of Medical Sciences, Yazd, Iran and PhD Student, Departments of Biostatistics and Epidemiology, School of public health, Shahid Sadoughi University of Medical Sciences, Yazd, Iran

DOI:

https://doi.org/10.13052/jrss0974-8024.19218

Keywords:

Linear classification, logistic regression, discriminant analysis, statistical learning

Abstract

Background: Classification is a core problem in statistical learning. Linear classification methods offer significant strengths, especially in educational and behavioral research. This tutorial introduces logistic regression (binary, multinomial, ordinal) and discriminant analysis methods.

Methods: Linear classification methods are categorized into three groups: discriminant analysis, logistic regression, and hybrid approaches. Each method’s mathematical foundations and implementation algorithms are explored, emphasizing their application in education and behavioral research.

Results: Logistic regression models are effective for handling binary, multi-class, and ordinal outcomes. Simulations in R illustrate their practical application across different contexts.

Conclusion: Linear classification methods are valuable for their interpretability, statistical efficiency, and theoretical rigor. Future research will focus on regularization, computational efficiency, and integration with nonlinear approaches.

Downloads

Download data is not yet available.

Author Biographies

Farzan Madadizadeh, Medical Informatics Research Center, Institute for Futures Studies in Health, Kerman University of Medical Sciences, Kerman, Iran and Center for healthcare Data modeling, Departments of Biostatistics and Epidemiology, School of public health, Shahid Sadoughi University of Medical Sciences, Yazd, Iran

Farzan Madadizadeh is an Associate Professor of Biostatistics at the Institute for Futures Studies in Health, Kerman University of Medical Sciences, Kerman, Iran. His research interests lie at the intersection of biostatistics and health informatics, with a particular focus on translating complex statistical methodologies into clear, practical knowledge for health researchers. He is actively involved in the design and psychometric evaluation of questionnaires, as well as the application of machine learning techniques to analyze healthrelated data and improve clinical decision-making. Dr. Madadizadeh has authored numerous peer-reviewed publications and is deeply committed to enhancing statistical literacy among students, clinicians, and public health professionals through accessible education and applied research.

Rashed Pourhamidi, PhD Student, Departments of Biostatistics and Epidemiology, School of public health, Shahid Sadoughi University of Medical Sciences, Yazd, Iran and Non Communicable Diseases Research Center, Bam University of Medical Sciences, Bam, Iran

Rashed Pourhamidi is a Ph.D. Student in Biostatistics at Shahid Sadoughi University of Medical Sciences, Yazd, Iran. His research interests include biostatistics, data science, and the application of statistical and computational methods in health research. He is particularly interested in analyzing health-related data and applying data-driven approaches to address methodological challenges in epidemiological and clinical studies. He has contributed to several peer-reviewed publications and has served as a reviewer for scientific journals. His work focuses on promoting rigorous research methods and evidence-based practice in health sciences.

Nahid khoddami, Center for healthcare Data modeling, Departments of Biostatistics and Epidemiology, School of public health, Shahid Sadoughi University of Medical Sciences, Yazd, Iran and PhD Student, Departments of Biostatistics and Epidemiology, School of public health, Shahid Sadoughi University of Medical Sciences, Yazd, Iran

Nahid khoddami is a Ph.D. Student in Biostatistics at Shahid Sadoughi University of Medical Sciences, Yazd, Iran. Her research focuses on developing and applying advanced statistical methods to solve complex data-driven problems in healthcare and medical sciences. She has a strong interest in longitudinal analysis and machine learning, aiming to improve decision-making processes through rigorous data analysis. Having authored multiple peer-reviewed articles, she is focused on bridging the gap between theoretical statistics and practical health research applications.

References

Abdillah, A., Sutisna, A., Tarjiah, I., Fitria, D., and Widiyarto, T. (2020). Application of multinomial logistic regression to analyze learning difficulties in statistics courses. Journal of Physics: Conference Series, 1490(1), 012012. https://doi.org/10.1088/1742-6596/1490/1/012012.

Abreu, M. N. S., Siqueira, A. L., Cardoso, C. S., and Caiaffa, W. T. (2008). Ordinal logistic regression models: Application in quality of life studies. Cadernos de Saúde Pública, 24, S581–S591. https://doi.org/10.1590/s0102-311x2008001600010.

Ananth, C. V., and Kleinbaum, D. G. (1997). Regression models for ordinal responses: A review of methods and applications. International Journal of Epidemiology, 26(6), 1323–1333. https://doi.org/10.1093/ije/26.6.1323.

Araveeporn, A. (2023). Comparison of logistic regression and discriminant analysis for classification of multicollinearity data. WSEAS Transactions on Mathematics, 22, 1–8. https://doi.org/10.37394/23206.2023.22.1.

Azeraf, E., Monfrini, E., and Pieczynski, W. (2021). Using the naive Bayes as a discriminative model. In Proceedings of the 13th International Conference on Machine Learning and Computing (pp. 106–110). Association for Computing Machinery. https://doi.org/10.1145/3457682.3457697.

Banerjee, D., Kumar, R., Tripathi, S., and Murry, B. (2024). Application of binary logistic regression in biological studies. Journal of the Practice of Cardiovascular Sciences, 10(1), 48–52. https://doi.org/10.4103/jpcs.jpcs_10_24.

Bhamre, N., Ekhande, P. P., and Pinsky, E. (2025). Enhancing naive Bayes algorithm with stable distributions for classification. Computer Science and Information Technology, 14(14), 107–114.

Boedeker, P., and Kearns, N. T. (2019). Linear discriminant analysis for prediction of group membership: A user-friendly primer. Advances in Methods and Practices in Psychological Science, 2(3), 250–263. https://doi.org/10.1177/2515245919849378.

Cichosz, P. (2015). Data mining algorithms: Explained using R. John Wiley and Sons.

Dutta, N., Subramaniam, U., and Padmanaban, S. (2019). Mathematical models of classification algorithm of machine learning. QScience Proceedings, 2019(1), 3. https://doi.org/10.5339/qproc.2019.epwlc2019.3.

Efron, B., and Hastie, T. (2021). Support-vector machines and kernel methods. In Computer age statistical inference: Algorithms, evidence, and data science (pp. 387–406). Cambridge University Press. https://doi.org/10.1017/9781108914062.024.

El-Habil, A. M. (2012). An application on multinomial logistic regression model. Pakistan Journal of Statistics and Operation Research, 8(2), 271–291. https://doi.org/10.18187/pjsor.v8i2.234.

Fagerland, M. W. (2014). adjcatlogit, ccrlogit, and ucrlogit: Fitting ordinal logistic regression models. The Stata Journal, 14(4), 947–964. https://doi.org/10.1177/1536867X1401400413.

Finch, W. H., and Schneider, M. K. (2006). Misclassification rates for four methods of group classification: Impact of predictor distribution, covariance inequality, effect size, sample size, and group size ratio. Educational and Psychological Measurement, 66(2), 240–257. https://doi.org/10.1177/0013164405282452.

Friedman, J. H., Hastie, T., and Tibshirani, R. (2010). Regularization paths for generalized linear models via coordinate descent. Journal of Statistical Software, 33(1), 1–22. https://doi.org/10.18637/jss.v033.i01.

Ghojogh, B., and Crowley, M. (2019). Linear and quadratic discriminant analysis: Tutorial. arXiv. https://arxiv.org/abs/1906.02590.

Górecki, T., and Smaga, Ł. (2019). fdANOVA: An R software package for analysis of variance for univariate and multivariate functional data. Computational Statistics, 34(2), 571–597. https://doi.org/10.1007/s00180-018-0844-4.

Gorriz, J. M., and Suckling, J. (2020). A connection between the pattern classification problem and the general linear model for statistical inference. arXiv. https://arxiv.org/abs/2012.08903.

Guo, S., and Tracey, H. (2020). Discriminant analysis for radar signal classification. IEEE Transactions on Aerospace and Electronic Systems, 56(4), 3134–3148. https://doi.org/10.1109/TAES.2020.2966141.

Haggag, M. M. M. (2018). Adjusting the penalized term for the regularized regression models. Afrika Statistika, 13(2), 1609–1630. https://doi.org/10.16929/as/1609.117.

Hastie, T., Tibshirani, R., and Buja, A. (1994). Flexible discriminant analysis by optimal scoring. Journal of the American Statistical Association, 89(428), 1255–1270. https://doi.org/10.1080/01621459.1994.10476866.

Hastie, T., Tibshirani, R., and Friedman, J. (2009). The elements of statistical learning (2nd ed.). Springer. https://doi.org/10.1007/978-0-387-84858-7.

Jin, R., Yan, F., and Zhu, J. (2015). Application of logistic regression model in an epidemiological study. Science Journal of Applied Mathematics and Statistics, 3(5), 225–229. https://doi.org/10.11648/j.sjams.20150305.13.

Krasoulis, A., Nazarpour, K., and Vijayakumar, S. (2017). Use of regularized discriminant analysis improves myoelectric hand movement classification. In Proceedings of the 8th International IEEE/EMBS Conference on Neural Engineering (NER) (pp. 395–398). IEEE. https://doi.org/10.1109/NER.2017.8008373.

Kravets, P., Pasichnyk, V., Prodaniuk, M., and Kis, Y. A. (2024). Computer modelling of logistic regression for binary classification. Visnyk Natsionalnoho Universytetu “Lvivska Politekhnika”. Seriia: Informatsiini Systemy ta Merezhi, 16, 167–190. https://doi.org/10.23939/sisn2024.16.167.

Kwak, C., and Clayton-Matthews, A. (2002). Multinomial logistic regression. Nursing Research, 51(6), 404–410. https://doi.org/10.1097/00006199-200211000-00009.

Lebret, R., Iovleff, S., Langrognet, F., Biernacki, C., Celeux, G., and Govaert, G. (2015). Rmixmod: The R package of the model-based unsupervised, supervised, and semi-supervised classification Mixmod library. Journal of Statistical Software, 67(6), 1–29. https://doi.org/10.18637/jss.v067.i06.

Lee, K., Ahn, H., Moon, H., Kodell, R. L., and Chen, J. J. (2013). Multinomial logistic regression ensembles. Journal of Biopharmaceutical Statistics, 23(3), 681–694. https://doi.org/10.1080/10543406.2012.756502.

Li, T., Prasad, A., and Ravikumar, P. K. (2015). Fast classification rates for high-dimensional Gaussian generative models. Advances in Neural Information Processing Systems, 28. https://dl.acm.org/doi/10.5555/2969239.2969357.

Madadizadeh, F., Soodejani, M. T., and Bahariniya, S. (2026). An Educational Tutorial on Fisher’s Exact Test for Medical Researchers. Journal of Reliability and Statistical Studies, 19(01), 199–214. https://doi.org/10.13052/jrss0974-8024.1919.

Martinez-Arroyo, M., and Sucar, L. E. (2006). Learning an optimal naive Bayes classifier. In Proceedings of the 18th International Conference on Pattern Recognition (ICPR’06) (Vol. 3, pp. 1236–1239). IEEE. https://doi.org/10.1109/ICPR.2006.741.

Morgan, G. A., Vaske, J. J., Gliner, J. A., and Harmon, R. J. (2003). Logistic regression and discriminant analysis: Use and interpretation. Journal of the American Academy of Child and Adolescent Psychiatry, 42(8), 994–997. https://doi.org/10.1097/01.CHI.0000046808.57140.71.

Nikita, E., and Nikitas, P. (2020). Sex estimation: A comparison of techniques based on binary logistic, probit and cumulative probit regression, linear and quadratic discriminant analysis, neural networks, and naive Bayes classification using ordinal variables. International Journal of Legal Medicine, 134(3), 1213–1225. https://doi.org/10.1007/s00414-020-02263-4.

Oliveira, L. L., Jiang, X., Babu, A. N., Karajagi, P., and Daneshkhah, A. (2024). Effective natural language processing algorithms for early alerts of gout flares from chief complaints. Forecasting, 6(1), 224–238. https://doi.org/10.3390/forecast6010013.

Peng, D., and Sha, J. (2021). Efficient HLS implementation of fast linear discriminant analysis classifier. IEEE Embedded Systems Letters, 13(4), 214–217. https://doi.org/10.1109/LES.2021.3078180.

Pohar, M., Blas, M., and Turk, S. (2004). Comparison of logistic regression and linear discriminant analysis: A simulation study. Metodološki Zvezki, 1(1), 143–160. https://doi.org/10.51936/ayrt6204.

Quintana, M., Guàrdia, J., Sánchez-Benavides, G., Aguilar, M., Molinuevo, J. L., Robles, A., Barquero, M. S., Antúnez, C., Martínez-Parra, C., and Frank-García, A. (2012). Using artificial neural networks in clinical neuropsychology: High performance in mild cognitive impairment and Alzheimer’s disease. Journal of Clinical and Experimental Neuropsychology, 34(2), 195–208. https://doi.org/10.1080/13803395.2011.630651.

Ramey, J. A., Stein, C. K., Young, P. D., and Young, D. M. (2016). High-dimensional regularized discriminant analysis. arXiv. https://doi.org/10.48550/arXiv.1602.01182.

Rouhizadeh, H., Yazdani, A., Zhang, B., and Teodoro, D. (2025). Exploring zero-shot cross-lingual biomedical concept normalization via large language models. medRxiv. https://doi.org/10.3233/SHTI250467

Saran, N. A., and Nar, F. (2025). Fast binary logistic regression. PeerJ Computer Science, 11, e2579. https://doi.org/10.7717/peerj-cs.2579.

Sharma, A., and Paliwal, K. K. (2008). Rotational linear discriminant analysis technique for dimensionality reduction. IEEE Transactions on Knowledge and Data Engineering, 20(10), 1336–1347. https://doi.org/10.1109/TKDE.2008.53.

Sifaou, H., Kammoun, A., and Alouini, M.-S. (2018). Improved LDA classifier based on spiked models. In Proceedings of the IEEE 19th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC) (pp. 1–5). IEEE. https://doi.org/10.1109/SPAWC.2018.8446025.

Tabassum, A., Sharma, M., Kumar, B., Dwivedi, S., Gupta, L. M., and Guleria, S. (2026). Prediction of Wheat Yield Through Soil Nutrient: Machine Learning and Feature Selection Approaches. Journal of Reliability and Statistical Studies, 19(01), 215–240. https://doi.org/10.13052/jrss0974-8024.19110.

Tharwat, A., Gaber, T., Ibrahim, A., and Hassanien, A. E. (2017). Linear discriminant analysis: A detailed tutorial. AI Communications, 30(2), 169–190. https://doi.org/10.3233/AIC-170729.

Tijani, A., Molyet, R., and Alam, M. (2022). Collision warning system using naive Bayes classifier. Technium, 4(5), 1–10. https://doi.org/10.47577/technium.v4i5.6653.

Wan, H., Wang, H., Guo, G., and Wei, X. (2017). Separability-oriented subclass discriminant analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(2), 409–422. https://doi.org/10.1109/tpami.2017.2672557.

Wurm, M. J., Rathouz, P. J., and Hanlon, B. M. (2021). Regularized ordinal regression and the ordinalNet R package. Journal of Statistical Software, 99(6), 1–42. https://doi.org/10.18637/jss.v099.i06.

Yu, H., and Yang, J. (2001). A direct LDA algorithm for high-dimensional data—with application to face recognition. Pattern Recognition, 34(10), 2067–2070. https://doi.org/10.1016/S0031-3203(00)00162-X.

Zhao, S., Zhang, B., Yang, J., Zhou, J., and Xu, Y. (2024). Linear discriminant analysis. Nature Reviews Methods Primers, 4(1), 70. https://doi.org/10.1038/s43586-024-00346-y.

Zhou, X., Chen, W., and Li, Y. (2024). netQDA: Local network-guided high-dimensional quadratic discriminant analysis. Mathematics, 12(23), 3823. https://doi.org/10.3390/math12233823.

Zorarpacı, E. (2021). A hybrid dimension reduction based linear discriminant analysis for classification of high-dimensional data. In Proceedings of the IEEE Congress on Evolutionary Computation (CEC) (pp. 1028–1036). IEEE. https://doi.org/10.1109/CEC45853.2021.9504951.

Zou, H., and Hastie, T. (2005). Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67(2), 301–320. https://doi.org/10.1111/j.1467-9868.2005.00503.x.

Downloads

Published

2026-09-25

How to Cite

Madadizadeh, F., Pourhamidi, R., & khoddami, N. (2026). A Tutorial on Linear Classification Methods for Educational and Behavioral Research with R. Journal of Reliability and Statistical Studies, 19(02), 689–716. https://doi.org/10.13052/jrss0974-8024.19218

Issue

Section

Articles