Data Science for Decision-Making in Public Health: A robust Clustering Approach for Large-Scale Weighted and Mixed-Type Socio-Economic Data
Published in Submitted for publication, 2027
This paper presents a new distance-based unsupervised statistical learning framework designed for large-scale weighted mixed-type socio-economic data to identify multidimensional health and well-being profiles among older Europeans. Using data from the Survey of Health, Ageing and Retirement in Europe (SHARE) across 26 countries, the pipeline combines composite indicators with robust generalized Gower distance and a fast $k$-medoids clustering algorithm for weighted data, uncovering five empirically distinct profiles—from highly resilient, socially integrated groups to multidimensionally vulnerable profiles—to guide efficient public resource allocation. Status: submitted for publication at Socio-Economic Planning Sciences.
Recommended citation: Albarrán, I., Grané, A., & Scielzo-Ortiz, F. (2026). "Data Science for Decision-Making in Public Health: A robust Clustering Approach for Large-Scale Weighted and Mixed-Type Socio-Economic Data." (Submitted for publication at Socio-Economic Planning Sciences).
