Psychological Testing: from good questions to defensible decisions
Complete exam-focused revision notes on test types, construction, item analysis, standardization, intelligence, personality, attitudes, computer-based testing and real-world applications.
किसी test का असली मूल्य उसके questions में नहीं, बल्कि इस बात में है कि उसके scores कितने विश्वसनीय, वैध और न्यायपूर्ण हैं।
✓ Full syllabus map⚡ NET/JRF traps🧪 4 interactive labs🧠 Bilingual support
🔁 Reliability
🎯 Validity
📏 Norms
⚖️ Fairness
No matching module found.
Try a broader term such as “validity”, “intelligence” or “attitude”.
01
Test Foundations — Know the Instrument Before the Score
What a psychological test is, and the major ways tests are classified
Psychological test: exam-ready definition
A psychological test is a standardized, objective and systematic procedure for obtaining a sample of behaviour and describing it with scores or categories.
यह पूरे व्यक्ति को सीधे नहीं मापता; यह व्यवहार का एक sample लेकर किसी construct के बारे में inference बनाता है।
Blueprint content validity की पहली सुरक्षा-दीवार है।
Item-writing rules
One clear idea per item; simple, age-appropriate language.
Avoid double negatives, clues, jargon and needless length.
MCQ stem should contain the problem; options must be plausible and grammatically parallel.
Avoid “always/never” unless logically necessary.
Check cultural, gender, disability and language bias.
Selected-response items
MCQ, true/false, matching.
Strength: objective and efficient scoring.
Risk: guessing and cueing.
VS
Constructed-response items
Short answer, essay, performance task.
Strength: samples organization and production.
Risk: scorer subjectivity; needs a rubric.
🧠 Construction chain: Define → Design → Draft → Debias → Dry run → Diagnose → Demonstrate quality → Document.
03
Item Analysis — Put Every Question Under the Microscope
Difficulty, discrimination, distractor quality and item bias
Difficulty index (p)
p = R / N
Proportion answering correctly. Higher p = easier item.
For dichotomous items, q = 1 − p.
Discrimination index (D)
D = pᵤ − pₗ
How strongly the item separates high scorers from low scorers.
Positive high D is desirable; negative D is a warning.
Item-total relation
Point-biserial correlation is used when item score is truly dichotomous (0/1) and total score is continuous.
Corrected item-total correlation excludes that item from the total.
🧪 Difficulty lab
p = 0.50 · Moderate
At p = .50, a dichotomous item has maximum response variance p(1−p), often useful for discrimination.
Indicator
Common interpretation
Decision clue
p > .80
Very easy
Keep only if essential or for confidence/warm-up.
p ≈ .30–.70
Moderate
Often most useful for norm-referenced discrimination.
p < .20
Very difficult
Check ambiguity, content mismatch or key error.
D ≥ .40
Very good
Usually retain.
D .20–.39
Moderate
Review or revise.
D < 0
Reverse discrimination
Investigate miskey, ambiguity or multidimensionality.
Distractor analysis
A functional distractor attracts some lower-performing examinees. A distractor chosen by almost nobody is non-functional and should be revised.
Good distractors are plausible, mutually exclusive and free of obvious clues.
ICC, IRT and DIF
An Item Characteristic Curve plots probability of a correct response against latent ability. IRT parameters commonly include difficulty (b), discrimination (a) and guessing (c).
DIF: equally able groups show different item-response probabilities. It signals possible bias and demands substantive review.
🎯 NET trap: In achievement items, “difficulty index” is actually an easiness proportion. The larger the p-value, the easier the item.
04
Standardization — Make Every Score Speak the Same Language
Uniform procedure, representative norms and transparent interpretation
Uniformity
Same instructions, time limits, materials, scoring and environmental conditions.
Norm sample
Large and representative of the population for whom interpretations will be made.
Manual
Purpose, population, administration, scoring, norms, reliability, validity, fairness and limitations.
True score is the expected average over infinitely many equivalent measurements—not a perfectly knowable score. Random error makes observed scores fluctuate.
Observed score को “सच्चाई” न मानें; वह true component और measurement error का मिश्रण है।
Standard Error of Measurement
SEM = SD × √(1 − rₓₓ)
Higher reliability → smaller SEM → narrower confidence interval around a score.
Approximate 95% interval: observed score ± 1.96 SEM.
Standardization is not just norms
A test can have norms yet be poorly standardized if administration varies. Standardization is the whole common procedure, while norms are the reference distribution.
⚠️ Never transport norms blindly. Language, time period, region, age, education and culture can change score meaning. Norms require periodic review.
05
Reliability — Can the Signal Survive Repetition?
Consistency, sources of error and the correct coefficient for each situation
Signal
Stable individual differences attributable to the construct.
+
Error
Fluctuation from time, items, raters, conditions or scoring.
The prophecy formula predicts reliability after changing test length by factor n.
What changes reliability?
More good, homogeneous items usually raise internal consistency.
Restricted score range can lower correlation estimates.
Unclear items, fatigue, changing conditions and subjective scoring add error.
Very heterogeneous constructs may legitimately show lower alpha.
⚠️ Alpha is not proof of unidimensionality. A high alpha can arise from many repetitive items. Factor structure and content must also be examined.
🧠 Match the error: Time → test–retest; Forms → alternate; Items → internal consistency; Judges → inter-rater.
06
Validity — Does the Interpretation Hit the Right Target?
Evidence for the meaning and use of test scores
Score inference
Modern idea of validity
Validity concerns how strongly evidence and theory support the interpretation of scores for a proposed use. Tests are not simply “valid forever”; a use and population matter.
Validity score के अर्थ और उसके उपयोग पर लागू होती है—केवल test के नाम पर नहीं।
Reliability is necessary but not sufficient for validity. An instrument can be consistently wrong.
Cattell; aims to reduce—not eliminate—cultural influence.
Wechsler logic
Current Wechsler batteries organize subtests into index scores. Common domains include verbal comprehension, perceptual/fluid reasoning, working memory and processing speed; exact indexes vary by edition.
Interpret the profile, not just FSIQ
Consider confidence intervals, index discrepancies, language, education, sensory/motor conditions, motivation and cultural opportunity. Large scatter may make a single global score less representative.
⚠️ “Culture-fair” means reduced cultural loading, not culture-free measurement. Familiarity with testing, schooling and socioeconomic opportunity still matter.
09
Creativity Testing — Count Possibilities, Not Just Correct Answers
Divergent production, scoring dimensions and process models
Divergent thinking
Generates multiple varied responses to an open problem. It is central to creativity assessment, but creativity also requires usefulness, context and evaluation.
Guilford emphasized divergent production within the Structure of Intellect model.
Convergent thinking
Narrows alternatives to one best/correct answer. Traditional intelligence and achievement items often emphasize it.
Creativity में “बहुत सारे उत्तर” और intelligence item में “सबसे सही उत्तर” का फर्क याद रखें।
💧FluencyNumber of relevant ideas
🌈FlexibilityNumber/variety of categories
✨OriginalityRarity or novelty
🧵ElaborationDetail and development
Major measures
TTCT (Torrance Tests of Creative Thinking): verbal and figural tasks; commonly score fluency, originality, elaboration, flexibility/other creative strengths depending on form.
Alternative Uses Task: generate unusual uses for a common object.
Good fit = abilities + interests + values + personality + opportunities + constraints. No single test should dictate a career. Results are hypotheses for exploration, integrated with interview, experience and local opportunity.
Eysenck: Psychoticism, Extraversion, Neuroticism plus Lie scale.
MBTI
Jung-inspired preference typology; popular in development settings, but categorical typing has important psychometric limitations.
Projective technique
Stimulus / response
Association
Rorschach
10 inkblot cards; “What might this be?”
Perceptual and personality patterns; requires standardized system and training.
TAT
Ambiguous pictures; stories about past, present, future
Murray’s needs and environmental press; motivational themes.
CAT
Pictures designed for children
Child-oriented apperception themes.
Rotter Incomplete Sentences Blank
Complete sentence stems
Adjustment, attitudes, conflicts and concerns.
Word association
First response to stimulus words
Associations and emotional complexes.
Draw-a-Person / HTP
Human/house-tree-person drawings
Hypothesis-generating; interpretation must be cautious.
🎯 Projective hypothesis: ambiguous, unstructured material allows aspects of the inner world to be expressed. Do not confuse this with objective scoring.
12
Neuropsychological Testing — Behaviour as a Window to Brain Systems
Batteries and focused tests for cognitive functions
Purpose
Neuropsychological assessment identifies patterns of strengths and weaknesses in attention, memory, language, visuospatial skill, motor function and executive control. It supports diagnosis, rehabilitation planning and monitoring—but is integrated with history, medical data and functional observation.
Test/battery
Main domain
High-yield clue
Halstead–Reitan Battery
Broad brain–behaviour functioning
Fixed-battery tradition; includes Category, Tactual Performance, Speech-Sounds Perception and other tests.
Based on Luria’s functional approach; standardized battery of many items/scales.
Bender–Gestalt
Visual–motor integration
Copy geometric designs; screening, not a stand-alone localization tool.
Wisconsin Card Sorting Test
Executive function
Set shifting, abstraction, feedback use; perseverative errors are important.
Stroop Test
Inhibitory control / selective attention
Interference between word reading and ink-colour naming.
Trail Making Test
Visual scanning, speed, flexibility
Part B adds alternating set demand.
Benton Visual Retention Test
Visual perception and memory
Reproduction/recognition of geometric designs.
Wechsler Memory Scale
Multiple memory systems
Memory profile, often paired with broader assessment.
Fixed vs flexible battery
Fixed: same comprehensive battery for all; comparability but long administration.
Flexible: tests selected for referral question; efficient but depends strongly on examiner expertise.
Performance can be altered by
Premorbid ability, education, culture/language, sensory or motor problems, medication, fatigue, pain, mood and effort. Performance validity measures may be needed.
⚠️ A neuropsychological score rarely maps neatly to one brain spot. Modern interpretation focuses on networks and patterns, not simplistic localization.
13
Attitude Scales — Turn Evaluation into a Measurable Continuum
Likert, semantic differential and Stapel scales with live demonstrations
Likert’s summated ratings
Respondents indicate degree of agreement with several favourable and unfavourable statements. Item scores are summed; negative items are reverse-scored.
Statement: “Psychological tests should be explained to examinees in plain language.”
Choose a response to see the coded score.
Osgood’s semantic differential
A concept is rated between bipolar adjective pairs, commonly on a 7-point continuum. Major meaning dimensions include Evaluation, Potency and Activity.
UnfairFair
WeakPowerful
PassiveActive
Stapel scale — not “Staples”
A single adjective is rated on a unipolar scale, typically from +5 to −5, without a zero category. It is useful when a natural opposite adjective is difficult to find.
🧠 LLS: Likert = Levels of agreement; Semantic = two Labels; Stapel = Single adjective.
14
Computer-Based Testing — Faster Delivery, New Sources of Error
CBT, adaptive testing, security, accessibility and equivalence
🧭EstimateStart ability estimate
🎯SelectChoose informative item
⌨️RespondRecord answer and time
🔄UpdateRevise ability estimate
🏁StopPrecision or item rule met
Advantages
Rapid scoring and reporting
Multimedia and precise timing
Randomized forms and automated routing
Adaptive tests can reduce items while maintaining precision
Threats
Digital divide and computer anxiety
Hardware/network variation
Identity, cheating and item exposure
Privacy, surveillance and data breaches
Quality safeguards
Mode-equivalence and usability studies
Accessibility and accommodations
Encryption, access control, audit trails
Human review of automated decisions
Computerized Adaptive Testing (CAT)
CAT uses an item bank—often calibrated with IRT—to choose items near the examinee’s current ability estimate. It is not simply a computer-delivered fixed test. It requires a large secure calibrated bank, content-balancing rules and exposure control.
⚠️ Algorithmic score ≠ automatic fairness. Validate the model, audit group performance, protect privacy and preserve a route for explanation and appeal.
15
Applications — One Instrument, Different Decisions
Clinical, organizational, educational, counseling, military and career settings
🏥 Clinical
Diagnostic clarification, severity, risk, treatment planning and outcome monitoring.
Combine: interview + tests + observation + records.
🏢 Organizational & business
Selection, placement, promotion, training needs, leadership and team development.
Demand: job analysis, predictive validity and adverse-impact monitoring.
🎓 Education
Achievement, readiness, learning needs, giftedness, disability support and programme evaluation.
Demand: age/grade norms and appropriate accommodations.
🤝 Counseling
Self-understanding, emotional concerns, strengths, decision-making and progress.
Demand: collaborative feedback, not labels.
🎖️ Military
Selection, classification, specialist placement, readiness and leadership potential.
Demand: high reliability, security, standardization and fairness.
🧭 Career guidance
Integrates aptitude, interests, values, personality, achievement and opportunity.
Demand: exploration of options, not a single deterministic verdict.
Selection vs classification
Selection: Who should enter?
Placement/classification: Where will the admitted person fit best?
Diagnosis: What pattern/problem is present?
Evaluation: Did an intervention work?
Decision rule
The higher the stakes, the stronger the required evidence. Use multiple sources, report uncertainty and check consequences for different groups.
High-stakes decision में एक score को अकेले “final truth” न बनाएं।
🎯 Application clue: Validity is use-specific. A test validated for counseling exploration is not automatically valid for employee rejection.
16
Ethics, Fairness & Final Retrieval Practice
Protect the person behind the score—then test your revision
Standard administration, dignity, accessibility, rapport without coaching, security and observation of relevant behaviour.
After testing
Accurate scoring, contextual interpretation, confidential records, understandable feedback and limits of inference.
Ethical principle
What it requires
Common violation
Competence
Use tests only with adequate training and within scope.
Untrained interpretation of complex instruments.
Informed consent
Explain purpose, procedure, foreseeable use and limits.
Hidden high-stakes use.
Confidentiality
Restrict access and disclose only with authority/lawful basis.
Sharing identifiable scores casually.
Test security
Protect items, scoring keys and copyrighted materials.
Publishing live secure items.
Fairness
Appropriate norms, language, accessibility and bias review.
Using irrelevant barriers or outdated norms.
Feedback
Explain results accurately, respectfully and with uncertainty.
Reducing a person to a label or number.
⚠️ Testing can dehumanize, label or invade privacy when used carelessly. Ethical assessment treats scores as evidence within context, never as the whole person.
⚡ NET/JRF Retrieval Check
Score: 0 / 8
1. Which coefficient best fits consistency of dichotomously scored items?
KR-20 estimates internal consistency for dichotomous items.
2. A test predicts training performance measured six months later. This is:
The criterion occurs in the future, so the evidence is predictive.
3. An item with p = .90 is usually described as:
p is the proportion correct. A larger p means an easier item.
4. Which attitude scale uses bipolar adjective pairs?
Semantic differential uses pairs such as good–bad or strong–weak.
5. Perseverative errors are especially associated with:
WCST assesses set shifting and cognitive flexibility; perseveration is a key index.
6. Which one is an interest typology?
RIASEC is Holland’s vocational interest/person–environment typology.
7. If reliability increases while SD stays constant, SEM will:
SEM = SD√(1−r). As r rises, the error term shrinks.
8. A computerized adaptive test mainly differs because it:
CAT updates ability after responses and chooses the next informative item.
🧠 Last-minute chain: Plan → Write → Pilot → Analyze → Standardize → Establish reliability → Accumulate validity evidence → Norm → Use ethically.
Support the work♡
Help us sustain serious learning.
Research, content development and hosting all carry real costs. If this work has supported your preparation, a voluntary contribution helps it keep growing.