Category: Research Methods

  • How to Design a Questionnaire for Your FYP: Scales, Pilot Testing and Validity

    How to Design a Questionnaire for Your FYP: Scales, Pilot Testing and Validity

    A questionnaire is not a list of questions you find interesting. Every item must trace back to an indicator in your theoretical framework, which traces back to a definition, which traces back to a research question. Build it in that order and it will survive validity testing. Build it by instinct and you will discover the problem only after your pilot data comes back.

    These seven steps take you from framework to a piloted instrument ready for main data collection.

    Step 1: List your constructs and their indicators

    Before writing a single item, write out each variable in your study and the indicators or dimensions it breaks into. Those indicators come from the literature you reviewed, not from what seems reasonable to you.

    For example, if you are measuring employee performance and your framework defines it through quantity of work, quality of work and timeliness, then you have three indicators and every performance item must belong to one of them.

    Expected output: a table with three columns — construct, indicator, source in the literature. Keep this table. It becomes an appendix and answers most instrument questions in the viva before they are asked.

    Step 2: Decide whether to adapt or build from scratch

    Adapting a published instrument is almost always the better choice for an FYP, and it is not a shortcut. A validated instrument has already been tested on real respondents, and using one lets you compare your results against previous studies.

    Rules when adapting:

    1. Cite the original source properly in your methodology chapter.
    2. State clearly what you changed and why, whether wording, context or number of items.
    3. Re-test validity and reliability on your own sample. Results reported for the original version do not automatically transfer to your adapted version or to a different population.

    Build from scratch only when nothing suitable exists for your construct, and expect to spend significantly longer on piloting if you do.

    Step 3: Write items that measure one thing each

    Most items that fail validity testing fail for reasons visible before any data is collected. Check every item against this list:

    • No double-barrelled items. “The system is fast and easy to use” asks two questions, and a respondent who finds it fast but confusing cannot answer honestly.
    • No leading wording. “How excellent was the service?” presumes the answer. Use neutral phrasing.
    • No jargon or technical terms your respondents may not share. Write for the actual population, not for your examiners.
    • No negatives inside a positive scale unless you intend reverse-coded items and remember to recode them before analysis. Forgetting to recode is a very common source of nonsensical reliability results.
    • Keep items short. Long sentences get skimmed, and skimmed items produce noisy data.
    • One time reference. Do not mix “usually” and “last week” across items measuring the same construct.

    Because many Malaysian students survey respondents who are answering in a second language, plain wording matters more here than general guidance suggests. If your respondents will read in Malay, translate carefully and have the translation checked by someone fluent in both languages, then state the translation procedure in your methodology chapter.

    Small group of chairs arranged for a pilot test session
    Pilot respondents should resemble your real sample but be excluded from the main study.

    Step 4: Choose your response scale deliberately

    The scale you choose determines what analysis is available to you later, so decide it with your analysis plan in mind rather than by habit.

    • Five-point or seven-point agreement scales are the standard for attitude and perception constructs. Seven points give slightly finer discrimination; five are easier for respondents and quicker to complete.
    • An even-numbered scale removes the neutral midpoint and forces a direction. Use it deliberately, and justify it, because removing the midpoint from respondents who genuinely have no view introduces its own distortion.
    • Frequency scales suit behaviour rather than attitude, but define the anchors concretely. “Often” means different things to different people; “three or more times per week” does not.
    • Keep the scale consistent within a construct. Mixing scale types inside one construct breaks the summed score.

    Remember the consequence for analysis: a single item on an agreement scale is ordinal data, while a summed multi-item scale is routinely treated as continuous. That distinction decides which tests you can run, as covered in the guide on which statistical test to use for your FYP data.

    Step 5: Structure and sequence the questionnaire

    Order affects completion rates and data quality.

    1. Cover statement. Who you are, the purpose, how long it takes, confidentiality, and voluntary participation. Many ethics committees require specific wording here.
    2. Screening items, if you need to confirm respondents meet your criteria.
    3. Main construct items, grouped by construct with a short heading for each section.
    4. Demographic items last. Placing them at the end reduces early drop-off, and respondents who have already invested effort are more likely to complete them.
    5. Closing thanks and contact details for questions.

    Keep the whole thing as short as your framework allows. Every additional item costs completion rate, and a shorter instrument fully completed beats a comprehensive one abandoned halfway.

    Step 6: Pilot test before the main collection

    The pilot is where design problems surface cheaply. Run it with respondents who resemble your real sample but who are excluded from the main study. A commonly used working figure in Malaysian faculties is around 30 pilot respondents, though this is convention rather than a national rule, so check your own guidelines.

    The pilot gives you three things:

    1. Item validity. Correlate each item score against the total score for its construct and identify items that do not behave like the rest.
    2. Reliability. Compute internal consistency for each construct, on the retained items only.
    3. Practical feedback. How long did it take? Which items did people ask about? Ambiguity shows up as questions.

    The order matters: test item validity first, remove items that fail, then recompute reliability on what remains. Computing reliability while retaining items you have already declared invalid is a frequent error and an easy one for an examiner to catch.

    Step 7: Fix what failed, and know when the problem is bigger

    If one or two items fail, remove or rewrite them and note it in your methodology chapter. If more than about a third of items fail, the problem is usually not the respondents. Look for these causes:

    • Items that were never properly derived from the construct definition.
    • Double-barrelled or ambiguous wording that different respondents read differently.
    • Reverse-coded items that were not recoded before analysis.
    • A construct that is genuinely two constructs, with items splitting into distinct groups.

    Rewrite and pilot again. Letting a construct proceed with only two surviving items will produce a question at the viva that is difficult to answer well.

    Writing all of this into your methodology chapter

    Everything above needs to appear in the instrument section of your methodology chapter: structure, source of items, scale, translation procedure if any, pilot results, and any items removed. The full chapter structure is set out in the guide on how to write the methodology chapter of a Malaysian FYP or thesis.

    How much justification is expected rises with the level of your degree, as explained in the comparison of FYP, thesis and dissertation in Malaysia.

    From instrument to a chapter that holds together

    A well-built questionnaire still has to be described in prose that stays consistent with your framework, your sampling and your analysis plan. That description is where instrument sections most often come apart, usually because an item was added or removed after the framework was written.

    Tesify helps you draft the instrument and methodology sections so that constructs, indicators, items and analysis stay aligned as the document grows, with you remaining the author and responsible for every methodological decision.

    Draft your instrument section in Tesify

    Frequently asked questions

    How many items should each construct have?

    Three to five items per construct is a common working range, giving enough coverage to compute reliability while keeping the instrument manageable. Fewer than three makes internal consistency difficult to establish meaningfully.

    Can I use an online form instead of paper?

    Yes, and it is standard practice. Online collection speeds distribution and removes transcription errors. Check whether your ethics approval covers online data collection and how the platform stores responses, since data protection expectations apply.

    What reliability value is acceptable?

    A coefficient of 0.70 is the most widely cited lower bound for acceptable internal consistency. It is a scholarly convention rather than an official rule, and some fields apply different thresholds. Cite the source you are following.

    Is a very high reliability value good?

    Not necessarily. A value close to 1.0 often signals that your items are near-duplicates of each other rather than that your instrument is excellent. It suggests redundancy: several items asking the same thing in slightly different words.

    Do I need to translate my questionnaire?

    It depends on your respondents. If they are more comfortable in Malay, translation improves data quality. Where you translate, describe the procedure in your methodology chapter and have the translation reviewed by someone fluent in both languages.

    Should pilot respondents be included in the main sample?

    No. They have already seen the instrument, which may influence how they respond. Exclude them and state that exclusion explicitly in your methodology chapter.

    What if my response rate is low?

    Distribute more, extend the collection period if your timeline allows, and follow up politely. If you still fall short of your calculated sample size, report the achieved number honestly, state the response rate, and discuss possible non-response bias in your limitations.

    Can I add questions after data collection has started?

    No. Changing the instrument mid-collection means your respondents did not all answer the same thing, which invalidates comparison across the dataset. If something essential is missing, stop, revise, and restart collection.

    How do I handle incomplete responses?

    Decide your rule before you look at the data and state it in the chapter, for example excluding any response missing more than a set proportion of items. Report how many were excluded and why. A rule chosen after seeing results invites the suspicion that it was chosen to produce them.

    Do I need permission to use a published instrument?

    Many published instruments are freely available for academic use, but some require permission from the author or publisher. Check the terms, and where required, request permission by email early. Always cite the original regardless of whether permission was needed.

  • Which Statistical Test Should I Use for My FYP Data?

    Which Statistical Test Should I Use for My FYP Data?

    Three questions decide your statistical test: what your research question is asking, what type of data your dependent variable is, and how many groups you are comparing. Answer those three in order and the choice narrows to one or two tests. Choosing the test first and forcing the question to fit is the error that sends most methodology chapters back for revision.

    This guide works through those three questions and gives you a decision table you can check your own study against.

    Question 1: What is your research question actually asking?

    Statistical tests answer four broad kinds of question, and almost every FYP falls into one of them:

    • Is there a difference between groups? Do male and female students differ in academic stress? Did the intervention group outperform the control group?
    • Is there a relationship between variables? Does study time relate to results? Does service quality relate to satisfaction?
    • Does one thing predict another? Do compensation and work environment predict employee performance?
    • Is there an association between categories? Is faculty membership associated with preferred learning mode?

    Write your research question out and underline the verb. Words like differ, compare and effect of an intervention point to difference tests. Words like relationship, association and correlate point to correlation. Words like predict, influence and determine point to regression.

    Question 2: What type of data is your dependent variable?

    This is the question students most often get wrong, and it controls everything downstream.

    • Nominal. Unordered categories: faculty, gender, employment status, yes or no answers.
    • Ordinal. Ordered categories without equal intervals: ranking, satisfaction levels, a single Likert item.
    • Interval or ratio (continuous). Numbers where the distance between values is meaningful: age, income, test score, or a summed scale score.

    One point causes endless confusion in FYP work: a single Likert item is ordinal, but a summed or averaged scale built from several Likert items measuring the same construct is routinely treated as continuous in the social sciences. That is why studies using validated multi-item instruments can legitimately run correlation and regression on scale scores. If you are analysing one item on its own, treat it as ordinal.

    Question 3: How many groups or variables are involved?

    For difference questions, count the groups and note whether they contain the same people:

    • Two independent groups. Different people in each, for example male and female students.
    • Two related measurements. The same people measured twice, for example before and after training.
    • Three or more independent groups. For example students from four faculties.

    For prediction questions, count your independent variables: one means simple regression, more than one means multiple regression.

    Laptop screen showing statistical output beside printed survey forms
    Answer the three questions before opening your software, not after.

    The decision table

    Your question Dependent variable Groups or predictors Usual test Non-parametric alternative
    Difference Continuous 2 independent groups Independent samples t-test Mann-Whitney U
    Difference Continuous 2 related measurements Paired samples t-test Wilcoxon signed-rank
    Difference Continuous 3+ independent groups One-way ANOVA Kruskal-Wallis
    Difference Continuous 3+ related measurements Repeated measures ANOVA Friedman
    Association Nominal 2 categorical variables Chi-square test of independence Fisher’s exact test for small counts
    Relationship Continuous 2 continuous variables Pearson correlation Spearman correlation
    Prediction Continuous 1 predictor Simple linear regression
    Prediction Continuous 2+ predictors Multiple linear regression
    Prediction Binary outcome 1 or more predictors Binary logistic regression

    When do you need the non-parametric alternative?

    The tests in the fourth column assume, among other things, that your continuous outcome is reasonably normally distributed within groups. When that assumption fails badly, you move to the alternative in the fifth column, which works on ranks instead of raw values.

    You would typically switch when a normality test on your data indicates a clear departure from normality, when your sample is small, or when your dependent variable is genuinely ordinal rather than continuous. The trade-off is real: non-parametric tests are more robust but generally have less power to detect an effect that exists.

    Two cautions worth knowing. First, normality tests become very sensitive in large samples and can flag trivial departures, so inspect a histogram rather than relying on the test alone. Second, the assumption concerns the distribution within groups, not the shape of your whole dataset lumped together.

    What else must you check before reporting?

    1. Independence of observations. Each participant contributes one response, unless you are deliberately using a paired or repeated-measures design.
    2. Homogeneity of variance for t-tests and ANOVA. Most software reports this alongside the test, and there are corrected versions to use when it fails.
    3. Linearity for correlation and regression. A scatterplot answers this in seconds and can reveal a strong curved relationship that a correlation coefficient would report as near zero.
    4. Multicollinearity for multiple regression. Predictors that are too highly correlated with each other make individual coefficients unstable and hard to interpret.
    5. Expected cell counts for chi-square. When expected counts fall too low, the test becomes unreliable and Fisher’s exact test is the usual remedy.

    Report these checks in your results chapter. Examiners look for them, and a study that reports a failed assumption and explains how it was handled reads as more competent than one that quietly ignores it.

    Two worked examples

    Example 1. “Is there a significant difference in academic stress between first-year and final-year students?” The question asks about difference. The dependent variable is a summed stress scale, so continuous. There are two independent groups. That gives an independent samples t-test, or Mann-Whitney U if the stress scores are badly skewed.

    Example 2. “Do service quality and price fairness influence customer loyalty?” The verb is influence, so this is prediction. The outcome is a continuous loyalty scale. There are two predictors. That gives multiple linear regression, with linearity and multicollinearity checked before the results are interpreted.

    The mistake that causes most revisions

    Running a test that does not match the research question. A study whose research question asks whether a relationship exists, but whose analysis chapter reports group comparisons, will be sent back regardless of how correctly the arithmetic was done.

    The fix is a single sentence in your methodology chapter that connects the two explicitly: “Because research question two asks whether service quality predicts customer loyalty, multiple linear regression was used, with loyalty as the dependent variable and service quality and price fairness as predictors.” That sentence tells the examiner you chose deliberately rather than by habit.

    Does the level of your degree change the test?

    No. The same tests are available at every level, and a well-executed t-test at doctoral level is not a weakness if it answers the question. What rises with the level is how thoroughly you justify the choice and how carefully you discuss its limitations. The differences between the levels are set out in the comparison of FYP, thesis and dissertation in Malaysia.

    Once the test is settled, it needs to be written into your methodology chapter alongside your design, population and instrument. The guide on how to write the methodology chapter covers that structure step by step.

    Writing up your analysis so it reads as a decision

    Choosing correctly is half the work. The other half is writing the analysis so a reader can follow why each step happened, from assumption checks through to interpretation, without the chapter contradicting itself.

    Tesify helps you draft and structure the methodology and results chapters so that the test you chose, the assumptions you checked, and the way you report the outcome stay consistent throughout. Tesify does not run your statistics; you do that in your own software and remain responsible for every number.

    Draft your methodology and results chapters in Tesify

    Frequently asked questions

    Can I use a t-test for three groups by running it three times?

    You should not. Running repeated pairwise tests inflates the chance of finding a significant difference by chance. Use one-way ANOVA for three or more groups, then a post-hoc test to identify which specific pairs differ.

    What is the difference between correlation and regression?

    Correlation measures the strength and direction of a relationship without assigning direction of influence. Regression models one variable as an outcome predicted by others and gives you an equation. If your hypothesis says one thing affects another, regression matches it better.

    Can I treat Likert data as continuous?

    A summed or averaged multi-item scale is commonly treated as continuous in social science research. A single Likert item is ordinal. State plainly in your methodology chapter which you are doing and cite the convention you are following.

    How do I know whether my data is normally distributed?

    Combine evidence rather than relying on one signal: inspect a histogram, look at skewness and kurtosis values, and run a normality test. In larger samples the formal test can flag departures too small to matter, so the visual check carries real weight.

    What if my results are not significant?

    Report them exactly as they are and discuss why. A non-significant result is a legitimate finding. What will damage you in the viva is being unable to explain it, or worse, adjusting data to produce significance, which is a serious breach of academic integrity.

    Do I need to report effect size?

    It is increasingly expected and always strengthens your discussion. Significance tells you whether an effect is likely to be real; effect size tells you whether it is large enough to matter practically. Most software reports it, sometimes only when you enable the option.

    Which software should I use?

    Use whatever your university licenses and your supervisor recognises. All mainstream statistical packages produce identical results for these standard tests because the underlying formulas are the same. Familiarity matters more than features when a deadline is close.

    My sample is only 30 respondents. Which tests can I use?

    Standard tests still run, but they have limited power to detect effects, and normality assumptions matter more at small sample sizes. Non-parametric alternatives are often the safer choice. State the sample size as a limitation and avoid overstating what your results establish.

    Can I use chi-square with continuous data?

    Not directly. Chi-square works on frequency counts in categories. Some students convert continuous data into categories to force it, which throws away information and is generally discouraged. Use a test suited to continuous data instead.

    What is the difference between one-tailed and two-tailed tests?

    A two-tailed test looks for a difference in either direction; a one-tailed test looks in a specified direction only. Use one-tailed only when theory genuinely justifies predicting the direction in advance, and state that justification. Two-tailed is the safer default and the usual expectation.