Transcription of Bias and equivalence in cross-cultural assessment: …
1 Original articleBias and equivalence in cross - cultural assessment : an overview>Fons van de Vijvera,*, Norbert K. TanzerbaDepartment of Psychology, Tilburg University, Box 90153 5000 LE, Tilburg, The NetherlandsbUniversity of Graz, AustriaAbstractIn every cross - cultural study, the question as to whether test scores obtained in different cultural populations can be interpreted in the sameway across these populations has to be dealt with. Bias and equivalence have become the common terms to refer to the issue. Taxonomy of bothbias and equivalence is presented. Bias can be engendered by the theoretical construct (construct bias), the method such as the form of testadministration (method bias), and the item content (item bias). equivalence refers to the measurement level at which scores can be comparedacross cultures. Three levels of equivalence are possible: the same construct is measured in each cultural group but the functional form of therelationship between scores obtained in various groups is unknown (structural equivalence ), scores have the same measurement unit acrosspopulations but have different origins (measurement unit equivalence ), and scores have the same measurement unit and origin in allpopulations (full scale equivalence ).
2 The most frequently encountered sources of bias and their remedies are described. 2004 Published by Elsevier sum Dans toute tude interculturelle, il faut r soudre la question de savoir si les scores au test obtenus dans des populations culturellementdiff rentes peuvent tre interpr t s de la m me mani re dans ces populations. Les termes de biais et d quivalence sont ceux devenus habituelsquand on envisage ce probl me. On propose une taxonomie tant du biais que de l quivalence. Le biais peut tre produit par le constructth orique (biais de construct), par la m thode, par ex., par la forme d administration du test (biais de m thode), et par le contenu d item (biaisa item). L quivalence se rapporte au niveau de mesure auquel les scores peuvent tre compar s dans les cultures. Trois niveaux d quivalencesont possibles: le m me construct est mesur dans chaque groupe culturel mais l aspect fonctionnel de la relation entre les scores obtenus dansles diff rents groupes est inconnu ( quivalence structurelle); les scores ont la m me unit de mesure dans les populations mais ont diff rentesorigines ( quivalence d unit de mesure); les scores ont la m me unit de mesure et la m me origine dans toutes les populations ( quivalenced chelle compl te).
3 Les sources de biais les plus fr quemment rencontr es sont d crites ainsi que les moyens d y rem dier. 2004 Published by Elsevier :Bias; equivalence ; Construct bias; Method bias; Item bias; OverviewMots cl s :Biais ; quivalence ; Biais de construit ; Biais de m thode ; Biais d item ; RevueThis article will discuss bias and equivalence in cross - cultural assessment . We will start with taxonomy of bias andequivalence ( de Vijver and Leung, 1997a,b). A lot ofcross- cultural research involves the application of instru-ments in various linguistic groups. Thus, the types of multi-lingual studies and their impact on bias and equivalence arediscussed in the second section. The third section describescommon sources of bias. The question of how to identify andto remove bias is discussed in the fourth section.
4 Finally,conclusions are Bias and equivalence : definitions and BiasSuppose that a geography test contains the item What isthe capital of Poland? This test is administered to pupils in alarge international educational achievement survey. The pro-portion of correct answers to the item will depend on, amongwother things, the pupils level of intellectual abilities, thequality of their geography education, and the distance of theircountry to Poland. Assuming that samples have been care-fully composed, the question will enable an adequate com-parison of the differences in knowledge of this particular>Premi re parution: Eur. Rev. Appl. Psychol. 47 (1997) 263.* Corresponding authorRevue europ enne de psychologie appliqu e 54 (2004) 119 2004 Published by Elsevier across all countries. However, suppose that the domainof the test is broader and that this item is used to assessgeographical knowledge.
5 Distance of the country to Polandwill now become a nuisance variable. Pupils from centralEurope are put at an advantage in comparison with pupilsfrom, say, Australia and USA. Such problems, known as bias,are common in cross - cultural assessment . More generally,bias occurs if score differences on the indicators of a particu-lar construct ( , percentage of students knowing that War-saw is Poland s capital) do not correspond to differences inthe underlying trait or ability ( , geography knowledge).Inferences based on biased scores are invalid and often do notgeneralize to other instruments measuring the same under-lying trait or ability. equivalence can be defined as the oppo-site of bias. However, historically, they have slightly differentroots and as a consequence, they have become and remainedassociated with different aspects of cross - cultural score com-parisons.
6 Bias has become the generic term for nuisancefactors in cross - cultural score comparisons whereas equiva-lence tends to be more associated with measurement levelissues in cross - cultural score comparisons. Both bias andequivalence are pivotal concepts in cross - cultural assess-ment. equivalence of measures (or lack of bias) is a prere-quisite for valid comparisons across cultural above example may well serve to illustrate an impor-tant characteristic of bias and equivalence : Both concepts donot refer to intrinsic properties of an instrument but to cha-racteristics of a cross - cultural comparison of that about bias always refer to applications of aninstrument in a particular cross - cultural comparison. An in-strument that reveals bias in a comparison of German andJapanese individuals may not show bias in a comparison ofGerman and Danish history of psychology has shown various examples ofsweeping generalizations about differences in abilities andtraits of cultural populations which, upon close scrutiny,were based on psychometrically poor measures.
7 In order toavoid making such sweeping statements which may attractmuch initial attention but which eventually do a disservice tothe field, the absence of bias ( , equivalence ) should bedemonstrated instead of simply assumed (Poortinga andMalpass, 1986).In order to facilitate the examination of bias, the followingtaxonomy may be useful. Three kinds of bias are distin-guished here (Van de Vijver and Leung, 1997a,b; Van deVijver and Poortinga, 1997). The first one is construct bias. Itoccurs if the construct measured is not identical across cul-tural groups. Western intelligence tests provide a good ex-ample. In most general intelligence tests, there is an empha-sis on reasoning, acquired knowledge, and memory. Socialaspects of intelligence are often less emphasized. However,there is ample empirical evidence that the latter aspects maybe more prominent in non-Western settings ( ,Super,1983).
8 The term intelligence as commonly applied in psy-chology does not do justice to its specific domain of applica-tion which is education. Binet s assignment, to design a testto detect children with learning problems which led to thedevelopment of intelligent tests as we know them, is stilldiscernible. The domain of the tests would be more appro-priately called scholastic intelligence .A second example of construct bias can be found in thework on filial piety ( , behaviors associated with being agood son or daughter;Ho, 1996). Compared to Westernsocieties, children in Chinese societies have more and diffe-rent obligations towards their parents. The difference may becaused by education and (1996)found inTurkey that help with household chores lost salience forparents with increased education. Similarly, the value ofchildren as old age security for the parents decreases with thelevel of income.
9 Therefore, a comparison of filial piety acrosscultural populations is susceptible to construct bias. Whenbased on a Western conception, the instrument will not coverall relevant aspects in a non-Western context. Analogously,an instrument based on a Chinese concept will contain be-haviors such as the readiness to take care of one s parentsfinancially in their old age which are only marginally relatedto the Western concept of filial piety. When based on acollectivist notion, the instrument will be over inclusive andwill contain various items that may well show little interper-sonal variation and induce a poor reliability of the instrumentin a Western question to be asked is how to deal with constructbias: is it possible to compare filial piety between individualsliving in Western and non-Western cultures? Probably, theeasiest solution is to specify the theoretical conceptualizationunderlying the measure.
10 If the set of relevant Western beha-viors is a subset of the non-Western set, then the comparisoncan be restricted to the Western set while acknowledging theincompleteness of the measure for the non-Western second type is method bias. The term method bias iscoined because it derives from aspects described in of empirical papers. Three types of method bias can beenvisaged. First, incomparability of samples on aspects otherthan the target variable can lead to method bias (sample bias).For instance, cultural groups often differ in educational back-ground and, when dealing with mental tests, these diffe-rences can confound real population differences on a targetvariable. Intergroup differences in motivation can be anothersource of method bias caused by sample incomparability. Forinstance, subjects who have been frequently exposed to psy-chological tests will show less motivation than subjects forwhom the instrument and/or the test situation has high bias also refers to problems deriving from instru-ment characteristics (instrument bias).