Variable scaling in cluster analysis of linguistic data

Moisl, H

doi:10.1515/cllt.2010.004

Variable scaling in cluster analysis of linguistic data

Downloads

Full text for this publication is not currently held within this repository. Alternative links are provided below where available.

Abstract

Where the variables selected for cluster analysis of linguistic data are measured on different numerical scales, those whose scales permit relatively larger values can have a greater influence on clustering than those whose scales restrict them to relatively smaller ones, and this can compromise the reliability of the analysis. The first part of this discussion describes the nature of that compromise. The second part argues that a widely used method for removing disparity of variable scale, Z-standardization, is unsatisfactory for cluster analysis because it eliminates differences in variability among variables, thereby distorting the intrinsic cluster structure of the unstandardized data, and instead proposes a standardization method based on variable means which preserves these differences. The proposed mean-based method is compared to several other alternatives to Z-standardization, and is found to be superior to them in cluster analysis applications.

Publication metadata

Author(s): Moisl H

Publication type: Article

Publication status: Published

Journal: Corpus Linguistics and Linguistic Theory

Year: 2010

Volume: 6

Issue: 1

Pages: 75-103

Print publication date: 14/06/2010

Date deposited: 15/07/2010

ISSN (print): 1613-7027

ISSN (electronic): 1613-7035

Publisher: Walter de Gruyter

URL: http://dx.doi.org/10.1515/cllt.2010.004

DOI: 10.1515/cllt.2010.004

Altmetrics

Altmetrics provided by Altmetric

ePrints

Variable scaling in cluster analysis of linguistic data

Downloads

Abstract

Publication metadata

Altmetrics

Share