How Do You Prove Validity and Reliability of a Translated Questionnaire in a Malaysian Nursing Thesis? CVI, Back-Translation and Cronbach’s Alpha (2026)

To prove validity and reliability of a translated questionnaire in a Malaysian nursing thesis, document the translation process (independent forward translations, a synthesis, back-translation and an expert committee), have a panel rate each item for relevance and report item-level and scale-level content validity indices, pilot the Malay version on a small group, and report Cronbach’s alpha from your own data. Then write all of it in Chapter 3 with the author’s original validation cited separately.

Why do examiners check validity and reliability so closely in nursing theses?

Nursing and health science projects measure knowledge, attitudes, anxiety, satisfaction, quality of life and practice. Most of the well-known scales were written in English for another country. If you hand a Malay version to nurses or patients without evidence that it measures the same thing, your results can be questioned no matter how good your statistics are. Examiners therefore look for two separate pieces of evidence: validity (does the instrument measure what it claims to measure in this population?) and reliability (does it measure consistently?).

The same logic applies if you use an English instrument with Malaysian participants who read mainly in Malay, or if you build a short questionnaire of your own. The only difference is how much of the process you must carry out yourself. Your faculty’s thesis handbook and your supervisor decide the exact expectations, so treat this guide as the method and the handbook as the rulebook. For the general mechanics of questionnaire building, see how to design a questionnaire for your FYP.

Do you need to translate and adapt the instrument, or can you use it as it is?

Start with a decision, because it determines the workload:

  1. A Malay version already exists with a published validation study. Use that version, cite the validation paper, get permission if required, and report reliability in your own sample. Our guide to validated Malay-language scales for a psychology FYP shows what an instrument section looks like when published validations exist, and the nursing thesis on older adults lists Malay-validated tools in a nursing context.
  2. Only the English version exists. You must translate and adapt it, and you should say so as a limitation if you cannot complete a full cross-cultural adaptation.
  3. You are building your own items. You need content validity from experts, a pilot test and reliability, but not back-translation unless you write items in one language first.

Always contact the instrument’s author or publisher about translation permission. A translation is a derivative work and some owners require approval.

What steps make up a defensible translation and cultural adaptation?

The reference most health researchers cite is Beaton, Bombardier, Guillemin and Ferraz (2000), “Guidelines for the process of cross-cultural adaptation of self-report measures”, Spine, 25(24), 3186 to 3191 (DOI 10.1097/00007632-200012150-00014). Studies that follow its method describe the same sequence of stages: forward translation, synthesis, back-translation, expert committee review, and pre-testing of the adapted version.

Stage What you do What you keep as evidence
1. Forward translation Two bilingual translators independently translate English to Malay. One knows the clinical concept, one does not. Both versions and translator profiles
2. Synthesis You and the translators reconcile the two versions into one. A table of differences and how each was resolved
3. Back-translation Two translators who have not seen the original translate the Malay version back into English. Both back-translations
4. Expert committee A panel compares all versions with the original and fixes meaning, wording, idioms and cultural fit. Minutes, the pre-final version, a log of changes
5. Pre-test A small group of the target population completes the pre-final version and explains what they understood. Comprehension notes and final wording changes
Forward and back translation workflow for adapting a questionnaire into Malay
Forward translation, synthesis, back-translation, committee review and pre-test form one traceable chain.

For a student project, the practical version is usually two forward translators, one reconciliation meeting with your supervisor, one back-translator and a language or subject expert reviewing the result. Describe precisely what you did and do not claim a full six-person process if you did a lighter one. Cultural adaptation matters as much as language: items about dining, religion, family structure or health-care access may need rewording for Malaysian participants without changing the construct.

How do you establish content validity with an expert panel?

Content validity asks whether the items cover the construct and are relevant to it. The method that nursing researchers use most is the content validity index (CVI). Lynn (1986), “Determination and quantification of content validity”, Nursing Research, 35(6), 382 to 386, is the classic reference for the procedure. Polit and Beck (2006), Research in Nursing & Health, 29(5), 489 to 497, analyse how the CVI is defined and calculated, and Polit, Beck and Owen (2007), Research in Nursing & Health, 30(4), 459 to 467, discuss how to interpret it.

  1. Choose experts. Recruit specialists in the clinical topic and, ideally, a language or measurement expert. Record their qualifications and years of practice.
  2. Rate relevance. Each expert rates every item on a four-point scale, for example 1 = not relevant, 2 = somewhat relevant, 3 = quite relevant, 4 = highly relevant.
  3. Compute the item-level CVI (I-CVI). The proportion of experts who rate the item 3 or 4.
  4. Compute the scale-level CVI (S-CVI). Polit and Beck (2006) show there are two methods: S-CVI/UA, the proportion of items rated relevant by all experts, and S-CVI/Ave, the average of the I-CVIs. They can give different values, so state which one you report.
  5. Apply a stated cut-off. Polit, Beck and Owen (2007) conclude that items with an I-CVI of .78 or higher for three or more experts can be considered evidence of good content validity. Cite the cut-off you use and keep it consistent.
  6. Revise or drop weak items, then record what changed.

A worked example, with invented numbers for illustration only: suppose seven experts rate a ten-item knowledge scale. Six items are rated 3 or 4 by all seven (I-CVI = 1.00), three items by six experts (6 divided by 7 = 0.857) and one item by five experts (5 divided by 7 = 0.714). The scale-level S-CVI/Ave is (6 x 1.00 + 3 x 0.857 + 0.714) divided by 10, which is 9.285 divided by 10, about 0.93. The S-CVI/UA is 6 divided by 10 = 0.60. The item with 0.714 falls below .78 and should be revised or removed. This is why reporting only one S-CVI can mislead a reader.

How do you run the pilot test and calculate reliability?

Reliability describes consistency. For a multi-item scale, the usual index is Cronbach’s alpha, introduced by Cronbach (1951) in Psychometrika, 16(3), 297 to 334. It reflects how strongly items within a scale vary together. For items scored right or wrong, such as a knowledge test, the KR-20 coefficient is the equivalent, and many statistics programs report it through the same reliability procedure. For stability over time you can use test-retest, with an interval long enough that participants do not simply remember their answers.

  1. Pilot the adapted questionnaire on respondents from your target population who are not in the main study.
  2. Enter the data and run the reliability analysis for each subscale separately, not only for the whole questionnaire.
  3. Read the item-total statistics and the “alpha if item deleted” column; investigate items that behave oddly instead of deleting them automatically.
  4. Report the alpha you obtained, and report it again for the main sample, because reliability belongs to the data you collected.

A common rule of thumb treats alpha of about .70 as the lower limit for acceptable internal consistency, but do not apply it blindly: alpha rises with the number of items, a very high value can signal redundant items, and a multidimensional scale needs a separate alpha for each dimension. Ask your supervisor which cut-off your faculty accepts. An illustrative result sentence: “Cronbach’s alpha was .86 for the 12-item practice subscale in the pilot (n = 30) and .88 in the main study (n = 210).” The numbers are invented; yours must come from your own output.

Expert panel rating questionnaire items for content validity in a nursing thesis
Each expert rates each item independently before the CVI is calculated.

What about construct validity, face validity and factor analysis?

Content validity and reliability are the minimum for most nursing FYPs. Two extras appear often:

  • Face validity. Whether the instrument looks appropriate and is understandable to respondents. Your pre-test comments and a short review by a few nurses provide this evidence.
  • Construct validity. Whether the items behave as the theory predicts, usually by factor analysis or by correlating with a related scale. Factor analysis needs a larger sample than a pilot provides, so most undergraduate projects state that it was not done and list it as a limitation or a suggestion for future work. For postgraduate theses, your supervisor may expect it.

Never claim a type of validity you did not test. “The instrument is valid and reliable because it was used in other studies” is the sentence examiners circle first.

What should you write in Chapter 3 about the instrument?

Use a short, fixed structure so the evidence is easy to find:

  1. Instrument name, authors, year, construct measured, number of items, subscales, response scale and scoring.
  2. Original validity and reliability evidence, cited to the source you opened.
  3. Translation and adaptation process, with the stages you actually did.
  4. Content validity: panel size, rating scale, I-CVI and S-CVI (and which method), cut-off, and the changes made.
  5. Pilot test: sample, setting, changes made after it, and the pilot reliability.
  6. Reliability in the main study, per subscale.
  7. Limitations of the adaptation.

Pair this section with a clear sample size rationale; how many respondents you need for an FYP covers that, and ethics approval for the pilot is covered in what ethics approval a nursing thesis needs in Malaysia. When the data arrive, which statistical test to use helps you match tests to variables.

What mistakes make examiners send back the instrument section?

  • Citing the original alpha only, with no alpha from the Malay version or your sample.
  • One translator, no back-translation, no description. The reader cannot tell what was done.
  • Expert panel of one or two. Too few experts make the CVI unreliable.
  • Dropping items silently after the pilot without recording why.
  • Reporting alpha for a mixed scale as a single number when the instrument has several dimensions.
  • No permission from the instrument owner for translation or use.

How can Tesify help with the Chapter 3 instrument section?

The instrument section is repetitive to format but unforgiving to get wrong. Tesify is used by 9,000+ students across 15,000+ chapters. You write the thesis yourself: it is 100% written by you, and Tesify helps you structure the chapter around the evidence you collected.

Frequently asked questions

What is the difference between validity and reliability in a nursing thesis?

Validity is whether the instrument measures what it claims to measure in your population. Reliability is whether it measures consistently. You need evidence for both, and they are reported separately.

How many experts do I need for a content validity index?

Polit, Beck and Owen (2007) discuss cut-offs for three or more experts. Many projects use between five and ten, but follow the number your supervisor and faculty accept and report their qualifications.

What is an acceptable I-CVI?

Polit, Beck and Owen (2007) conclude that an I-CVI of .78 or higher for three or more experts can be considered evidence of good content validity. State the cut-off you choose and apply it to every item.

What is the difference between S-CVI/UA and S-CVI/Ave?

S-CVI/UA is the proportion of items that every expert rates relevant. S-CVI/Ave is the average of the item-level indices. Polit and Beck (2006) show that the two can differ, so state which one you report.

Is back-translation compulsory?

Back-translation is part of the widely cited cross-cultural adaptation process of Beaton and colleagues (2000). Your faculty may accept a lighter process for an FYP, but you must describe exactly what you did.

Do I need permission to translate a questionnaire?

Usually yes. Contact the author or publisher, keep the reply, and cite it in Chapter 3. Some instruments have an official translation process owned by the publisher.

What Cronbach’s alpha is acceptable?

A value around .70 is a common rule of thumb, but acceptable values depend on the number of items, the purpose and your faculty. Report the value for each subscale and discuss it honestly.

Can I use KR-20 instead of Cronbach’s alpha?

Yes, for items scored as right or wrong, such as a knowledge test. KR-20 is the dichotomous-item counterpart of alpha.

Should I report the original author’s reliability or my own?

Both, clearly separated. The original figure belongs to the original sample, while your pilot and main-study figures describe your Malay version and participants.

Can I skip the pilot test?

Check with your supervisor. A pilot is the normal way to find unclear items and to estimate reliability before the main study, and examiners expect to see one reported or a reason why it was not done.